Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2586 results about "Vision based" patented technology

Uncoupling robot control system and method based on multi-source visual fusion

The embodiment of the invention provides an unhooking robot control method based on multi-source visual fusion, which is applied to the technical field of robot control and comprises the following steps: acquiring an RGB image, a depth image, an infrared image and IMU data through a multi-source sensing system mounted at the tail end of a robot; carrying out feature fusion identification by adopting a double-branch neural network, and outputting the boundary contour of the lifting hook and the three-dimensional coordinates of the optimal grabbing point; the visual coordinates are unified to a robot base coordinate system through a registration correction mechanism; a Transform prediction model is constructed based on the visual and inertial signals, and future pose changes of the lifting hook are estimated; a feedforward control track is generated to counteract swing of the lifting hook, and track correction is carried out in combination with visual servo feedback; and a joint instruction is generated through path planning and inverse kinematics solution, and the mechanical arm is driven to complete precise unhooking operation. According to the method, the recognition precision, the anti-interference capability and the operation success rate of unhooking operation in complex illumination and dynamic environments are effectively improved.
Owner:ANHUI HUADIAN SUZHOU POWER GENERATION

Industrial part alignment method and system based on visual analysis and storage medium

The invention relates to the technical field of image processing, and discloses an industrial part alignment method and system based on visual analysis and a storage medium. The method comprises the steps that a three-view camera collects an industrial part image, and preprocessing is carried out through gradient magnitude local contrast enhancement to obtain an enhanced image; performing hierarchical feature extraction to identify edge contours and key control points to form a multi-dimensional feature set; and establishing a dynamic reference coordinate system based on the feature set to obtain a part space attitude matrix. And the attitude deviation is compensated through Z-axis offset and rotation coupling error analysis. Posture adjustment is decomposed into a plurality of sub-stages, an alignment track is optimized by adopting a variable speed planning strategy, and accurate alignment of the parts is achieved. The problems that multi-view visual information fusion is insufficient, a special recognition algorithm for geometrical characteristics of the industrial parts is lacked, and Z-axis offset and rotation coupling error compensation is inaccurate in the posture adjustment process are solved, and the precision and stability of alignment of the industrial parts are improved.
Owner:BEIJING TIANYUAN 3D TECH CO LTD

Dynamic Latent Space Adaptation Based on Spatiotemporal Kernal Context for Multiscale Rendering

A system for dynamic latent space adaptation using spatiotemporal kernel context for multiscale rendering with hierarchical and Lorentzian autoencoders. The Spatiotemporal Kernel Estimator (SKE) analyzes media through motion field, temporal recurrence, frequency band, and scene semantics analyzers to generate adaptive kernel parameters encoding content-specific importance distributions. The system dynamically adapts latent manifold geometry by modifying metric tensor properties according to kernel context, enabling content-aware compression that allocates representational capacity based on visual significance. A multiscale cache implements kernel-adaptive retention policies prioritizing important regions. An adaptive renderer provides intelligent level-of-detail selection based on zoom level and kernel-estimated importance, optimizing processing allocation. The self-optimizing architecture continuously refines kernel context and geometric adaptation based on user interaction and performance feedback, achieving superior compression ratios and perceptual quality. Applications include bandwidth-efficient video streaming, virtual reality, scientific visualization, and cognitive video analytics requiring intelligent context-aware visual processing.
Owner:ATOMBEAM TECH INC

Vision-language interaction-based remote sensing image open vocabulary segmentation method and system

The invention discloses a remote sensing image open vocabulary segmentation method and system based on vision-language interaction. According to the method, the advantages of a multi-modal large language model and a semantic segmentation network based on a visual basic model are cooperatively utilized, language-pixel two-way mapping is taken as a core, five types of marks of images, texts, categories, objects and segmentation are introduced as carriers of cross-modal information, and three types of cross-modal fine-grained information interaction modules are taken as bridges, so that the cross-modal information interaction is realized. Bidirectional mapping and alignment of fine-grained information of the multi-modal large language model and the semantic segmentation network are realized, the open vocabulary segmentation capability of the semantic segmentation model is improved, and the method can adapt to different remote sensing scenes and category definitions. The method has the following advantages: the method has high performance, strong generalization ability and good expansibility, can provide any category of semantic segmentation maps for unlabeled target domain images based on instructions, and has high application value in the aspects of urban planning, map making, disaster response and the like.
Owner:WUHAN UNIV

PCB defect detection method based on visual converter combined with conditional diffusion

The invention belongs to the technical field of computer vision and deep learning, and particularly relates to a PCB defect detection method based on combination of a visual converter and conditional diffusion. Comprising the following steps: constructing an unlabeled PCB image data set and carrying out data preprocessing and enhancement to obtain a preprocessed image; executing a self-supervised pre-training task on the preprocessed image to obtain a feature extraction network; based on a conditional diffusion model, generating a synthetic defect PCB image and a label thereof by using the features output by the feature extraction network and the defect type control vector; mixing the synthetic defect image with a small number of real defect images to construct a training set; performing training adjustment on the defect detection model by adopting the training set to obtain a trained defect detection model; performing PCB defect detection by using the trained defect detection model; according to the method, the robustness and the cross-domain generalization ability are remarkably improved, the missed detection risk is reduced, and the rapid and stable quality control requirement of the production line is met.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Semi-supervised image semantic segmentation method and system based on visual basic model

The invention provides a semi-supervised image semantic segmentation method and system based on a visual basic model, and the method comprises the steps: constructing a multi-task model which comprises a visual basic model and a depth estimation basic model, and the visual basic model is connected with a task solution head, an adapter parameter efficient fine tuning module and a multi-modal cross fusion module; the task solution head comprises a semantic segmentation head and a depth estimation head; extracting semantic hierarchy features and a depth feature map of the RGB image, performing cross attention fusion on the semantic hierarchy features and the depth feature map, and inputting obtained fusion features into a semantic segmentation head and a depth estimation head respectively; semi-supervised learning is adopted to train a multi-task model, only parameters in the adapter parameter efficient fine tuning module and the multi-modal cross fusion module are trained, and a multi-task loss function is adopted. The image semantic segmentation model obtained through training can improve semantic segmentation performance, reduce training cost and is suitable for different tasks.
Owner:SHANGHAI JIAOTONG UNIV

Tunnel illumination control method considering visual adaptability of driver

The invention relates to a tunnel illumination control method considering visual adaptability of a driver, and belongs to the technical field of tunnel illumination control, and the method comprises the following steps: S1, employing a volume scattering point location method to obtain visual adaptability parameters of visual adaptation states of all segments in a tunnel; s2, dynamically constructing a predictive visual adaptation model of the driver group based on the visual adaptation parameters, and outputting expected visual states of the drivers according to the predictive visual adaptation model; s3, according to the expected visual state, a parameter control mechanism is adopted to control target illumination parameters of all the sections of the tunnel; s4, according to the target lighting parameters, the working state of lighting equipment in a corresponding section in the tunnel is regulated and controlled, so that the light environment in the tunnel is matched with the dynamic visual adaptability of a driver, and safe visual guidance is achieved; the method has the beneficial effects that the target illumination parameters of all the sections of the tunnel are controlled by adopting a parameter control mechanism according to the expected visual state according to the target illumination parameters.
Owner:SICHUAN HIGHWAY ENG CONSULTING & SUPERVISION CO LTD

Robot control system and method based on visual sense and force sense fusion and robot

The invention discloses a robot control system and method based on vision and force sense fusion and a robot, and relates to the technical field of robots. The robot control system based on vision and force sense fusion and the control method thereof analyze visual features in real time through an artificial intelligence processing unit to predict mechanical parameters; and the self-adaptive impedance control unit is combined to dynamically fuse a position instruction and force sense feedback, so that precise adaptation of target object characteristics and compliant regulation and control of an interaction process are realized, real-time sensing and dynamic adjustment capabilities are achieved, assembly damage caused by mechanical mismatching is effectively avoided, and the reliability and efficiency of precise assembly are improved.
Owner:SHENZHEN HUACHENG IND CONTROL

Text and image bidirectional alignment method and system based on multi-hop parallel reasoning

The invention discloses a text and image bidirectional alignment method and system based on multi-hop parallel reasoning, and the method comprises the steps: analyzing and marking a text, and obtaining a multi-granularity text feature group; image generation is processed, grids and target output are combined into a multi-scale visual feature group, an alignment index table is established, a sub-target chain is constructed based on visual features, the position and the dependency relationship are recorded, and explicit constraints are injected into five elements; recalling the candidate area, generating an anchor point through verification and fusion, writing the anchor point into a cross-modal index, locally completing three-layer decoupling scoring based on anchor point geometry and the cross-modal index, and outputting a three-state evidence in combination with a threshold value; the method comprises the steps of generating an evidence set through geometric and semantic consistency calibration, generating cross-hop evidence according to a rule, weighting and pooling the evidence set, generating a shared vector and traceable meta-information, inputting a task head output result and generating an evidence list and an auditing path, and according to the method, local alignment, constraint verification and evidence accumulation are completed hop by hop. And high-precision, traceable and interpretable cross-modal correspondence is realized.
Owner:SHAANXI NORMAL UNIV

Prefabricated cabin welding seam quality detection method based on self-supervised learning

The invention discloses a prefabricated cabin welding seam quality detection method based on self-supervised learning, and the method comprises the following steps: collecting a multi-source welding seam image and process parameters, carrying out the synchronous calibration, and constructing a multi-modal observation matrix; the image is input into a DINOv2 model based on a visual Transform, dense features are obtained, and inter-frame registration and sequence reconstruction are completed; performing difference analysis on adjacent frames, extracting spatial changes and marking potential defects; the dense features and the process parameters of the corresponding time periods are fused, joint features are constructed and classified according to rules, and a preliminary result is output; inputting the defect area and the process parameters into an improved DEER model, extracting an influence path and amplitude, and generating a sensitivity score and a confidence coefficient; final judgment is given in combination with the preliminary result, the sensitivity and the confidence coefficient, and feedback is formed according to judgment and process difference. According to the method, high-precision and explainable detection and process optimization of weld defects are realized, and the method is suitable for online / offline quality control and tracing.
Owner:ANHUI HUANYU INTELLIGENT EQUIPMENT CO LTD

Robot dynamic grabbing control method based on visual sense and tactile sense depth fusion

The invention relates to a robot dynamic grabbing control method based on visual sense and tactile sense depth fusion. The method comprises the following steps that firstly, based on visual information, surface geometric features of a target object are extracted, grabbing adaptability scores are calculated, an optimal grabbing point is selected, and a collision-free approaching track is planned; secondly, in a contact establishment stage, detecting a contact event through a touch sensor, fusing vision and touch data to unify a coordinate system, calculating a vision-touch consistency score, and applying an initial grabbing force; and thirdly, in the stable holding stage, slip detection is conducted through wavelet packet energy entropy and pressure gradient, and the grabbing force and impedance parameters are dynamically adjusted in combination with self-adaptive impedance control. According to the method, intelligent control over the whole process from approaching to contact to stable holding is achieved, and the adaptability, stability and safety of robot grabbing in the dynamic environment are improved.
Owner:INEXBOT

Picking mechanical arm operation method and system based on visual feedback

The invention provides a picking mechanical arm operation method and system based on visual feedback. The picking mechanical arm operation method comprises the steps that firstly, a multi-dimensional visual perception set including the shape contour, surface physical characteristic visual representation, connection part space distribution, growth environment dynamic interference and the like of a picked object is obtained; a mechanical arm collaborative decision-making model is constructed based on the multi-dimensional visual perception set, stress distribution characteristics of connection parts are analyzed, and optimal contact area parameters and mechanical arm posture collaborative parameters are determined; according to the parameters, generating a sub-time sequence action instruction set containing action information in the contact stage, the clamping stage and the separation stage; when the mechanical arm is controlled to execute the instruction set, a dynamic visual feedback set of the picking scene is obtained in real time; and finally, input parameters of the collaborative decision model are updated based on feedback, and instruction set parameters are adjusted, so that the mechanical arm completes picking operation on the premise of keeping the integrity of the picked object, and the intelligent level and the picking operation quality of the picking mechanical arm are improved.
Owner:SICHUAN QIANXIAOMO TECH CO LTD

Visual large model-based scene reconstruction and semantic understanding method and system

The invention provides a scene reconstruction and semantic understanding method and system based on a visual large model, and relates to the field of computer vision and three-dimensional reconstruction, and the method comprises the steps: obtaining multi-view image data of a target scene, inputting a pre-training visual large model, and outputting a joint feature representation and attention weight matrix; generating three-dimensional coordinate values and semantic probability distribution of spatial sampling points in a three-dimensional space according to the joint feature representation, and constructing a spatial semantic field; clustering the spatial sampling points by using the attention weight matrix, performing semantic consistency enhancement, and converting the spatial sampling points into deterministic semantic tags; constructing a geometric optimization objective function, and adjusting three-dimensional coordinate values; and extracting continuous space sampling points with the same semantic tag to form an object boundary, constructing a scene topological graph, deducing a scene functional structure, and generating a navigation path. According to the method, the unification of accurate geometric reconstruction and deep semantic understanding of the scene is realized, the three-dimensional reconstruction precision and semantic analysis accuracy are improved, and reliable support is provided for intelligent navigation.
Owner:SMIC WANYE TECHNOLOGY CO LTD

Navigation instruction generation method, device and system based on multi-modal environment understanding

The invention provides a navigation instruction generation method, device and system based on multi-modal environment understanding, and the method comprises the steps: obtaining multi-modal perception information comprising the three-dimensional image data and three-dimensional point cloud data of an environment where a current unmanned aerial vehicle is located, and the task text information of the unmanned aerial vehicle; and performing semantic fusion extraction on the multi-modal perception information and the task text information based on a visual language fusion model to obtain a multi-modal embedded representation, constructing a large language model Prompt based on a plurality of prior navigation templates and the multi-modal embedded representation, and inputting the large language model Prompt into the large language model. According to the method, the navigation instruction text output by the large language model is obtained, then the navigation instruction text is analyzed into the control instruction sequence which can be recognized and executed by the unmanned aerial vehicle, the control instruction sequence is issued to the unmanned aerial vehicle, and the accuracy, continuity and stability of navigation can be maintained in a complex, dynamic and GNSS limited environment.
Owner:BEIJING SHENGSHI TIANAN TECH CO LTD

Multi-scene-oriented intelligent light adaptive control method and system

The invention provides a multi-scene-oriented intelligent light adaptive control method and system, and the method comprises the steps: recognizing the current behavior state of a user based on visual sensor data, predicting the behavior state of the user, and outputting a light control behavior tag; combining the behavior label with a scene type and ambient light intensity to generate a target illumination intention vector; according to the target illumination intention vector, in combination with an illumination control conflict situation, real-time lamp parameters and user space state data, calculating an illumination weight of each lamp to a user, distributing brightness and color temperature parameters, and in combination with the target dimming process duration, generating a lamp control parameter set; and converting the lamp control parameter set into a dimming instruction sequence, performing time synchronization and delay compensation, and issuing the dimming instruction sequence to each lamp for execution. According to the method, accurate mapping from behavior intention to physical illumination is realized, and the method has extremely high universality, deployability and engineering controllability.
Owner:GUANGZHOU NIGHTRAINBOW TECH

Vision-based drainage pipeline defect three-dimensional quantification method and device

The invention relates to the technical field of image processing, in particular to a drainage pipeline defect three-dimensional quantification method and device based on vision, and the method comprises the steps: carrying out the at least one processing of sparse point cloud reconstruction, dense point cloud reconstruction, grid reconstruction, texture mapping and scale calibration of a collected visual image, obtaining a real three-dimensional model of the drainage pipeline; calculating a central point of the drainage pipeline based on the real three-dimensional model, and calculating a central line of the drainage pipeline according to the central point; and based on the center line, identifying at least one actual defect of a staggered joint defect, a disjunction defect, a fluctuation defect and a deformation defect of the drainage pipeline, so as to output actual defect information of the drainage pipeline according to the at least one actual defect. Therefore, the problems of incomplete identification or misjudgment of information such as cracks, leakage and sediments and the like caused by low efficiency and detail missing due to relatively great influence of manual subjective judgment or sparse laser point cloud in related technologies are solved.
Owner:TSINGHUA UNIVERSITY

Construction method and system for real-time teleoperation of dual arm and hand robot based on vision guidance

The present invention relates to the field of robotic arm teleoperation technology, and specifically to a construction method and system for real-time teleoperation of a dual arm and hand robot based on vision guidance, which includes setting up a first experimental platform, acquiring pose information of the dual arm and hand robot and a homogeneous transformation matrix of an end joint and transforming a first coordinate system; setting up a second experimental platform; performing calibration by a plurality of cameras and transforming a second coordinate system; obtaining spatial transformation pose information of the end joint of the robotic arm under a tracking state, reading finger pose data, and performing real-time following. The real-time teleoperation system of the present disclosure includes a control system host, an operation glove, an optical locator, a positioning coordinate plate, a marker ball tool, and a dual arm and hand robot. The present disclosure introduces a method of identifying the marker ball tool by the optical locator into the teleoperation control of the dual arm and hand robot, which is more accurate and faster. Meanwhile, the method of remotely controlling the robotic arm by an operator wearing the operation gloves can provide better human-machine interaction function.
Owner:YANSHAN UNIV

Real-time back break control method of heading machine based on visual association and digital twinning

The invention discloses a real-time back break control method of a heading machine based on visual association and digital twinning, which relates to the technical field of intelligent control of coal mining equipment, and comprises the following four steps: processing binocular visual data through a double-domain double-branch Transformer model, extracting coal rock and roadway structure characteristics and outputting cutting head offset; a GCN-LSTM model is adopted to fuse visual, laser radar and fiber-optic gyroscope data, and dynamic calibration and output of a three-dimensional pose are carried out; a roadway three-dimensional model is constructed based on the SLAM technology, and real-time updating of a digital twin model is achieved through edge-cloud collaboration; and in combination with a reinforcement learning strategy, generating a cutting parameter adjustment instruction according to the deviation of the digital twin model, and correcting model parameters through error feedback. The method solves the problems of low visual recognition precision, sensor data drift, digital twinning update lag and the like in an underground complex environment, improves the back break control precision, tunneling efficiency and equipment stability, and is suitable for tunneling operation under complex geological conditions.
Owner:TAIYUAN INST OF CHINA COAL TECH & ENG GROUP +1

Robot unstacking and stacking support system based on 3D visual system

The invention discloses a robot unstacking and stacking support system based on a 3D visual system, and belongs to the technical field of 3D visual systems. Comprising a 3D visual data acquisition preprocessing module, a target positioning priority analysis module, a box type intelligent identification matching module, a dynamic path planning obstacle avoidance module, a self-adaptive stacking strategy generation module, an information binding tracing management module and a feedback optimization module. The problems of low efficiency and conflict caused by a single grabbing sequence in traditional stacking and unstacking operation are solved, the priority weight is dynamically adjusted by combining the inclination angle of the carton, weight distribution and reachability verification, and the optimality and safety of the grabbing action of the robot are ensured.
Owner:SUZHOU BOXTONG TECHNOLOGY CO LTD

Intelligent patch board spot welding track control method and system based on visual perception

The invention relates to the technical field of image data processing and intelligent control, and discloses a patch board spot welding track intelligent control method and system based on visual perception. According to the method, synchronous image streams are collected through a binocular vision sensor, and epipolar correction image pairs are generated through timestamp synchronization, ROI extraction and epipolar correction processing; calculating a disparity map by adopting a stereo matching algorithm, and constructing a compensated three-dimensional point cloud model based on deformation vector field fusion multi-frame point cloud data of a radial basis kernel function; welding spot position error vectors are generated by extracting welding spot feature points and performing spatial filtering optimization; and carrying out inverse kinematics calculation on the mechanical arm by adopting a damping least square method, carrying out safety constraint optimization in combination with prospective collision risk assessment and real-time pose data, and generating a trajectory compensation instruction. According to the method, the problems of dynamic deformation compensation and motion safety in patch plate welding are solved, the submillimeter welding spot positioning precision is achieved, and the welding quality stability and the system robustness are improved.
Owner:重庆衍数自动化设备有限公司

Reinforced learning training method and system for relieving hallusion of multi-modal large model

The invention discloses a reinforcement learning training method and system for relieving illusion of a multi-modal large model, and belongs to the field of reinforcement learning training of a multi-modal large language model. Firstly, a planning and visual description generation step is introduced in an early stage to guide a model to perform structured reasoning, then a grouping relative strategy optimization algorithm is used, reward values are calculated for multiple candidate responses generated by the model after cold start, and particularly, a visual perception reward mechanism is set. The reward mechanism evaluates the consistency of the generated text description and the visual information by using an external large language model. Then, based on a vision description attention score advantage distribution method, learning of the model on key vision signals is dynamically enhanced, and the perception ability of the model on the vision signals is improved; and finally, the perception and reasoning performance of the model is further improved by adopting multiple rounds of rejection sampling and supervised fine tuning. The scheme does not depend on a model architecture, the extra overhead is small, the illusion problem caused by early image-text inconsistency is effectively solved, and the accuracy and the reliability are improved.
Owner:ZHEJIANG UNIV +1

Unmanned aerial vehicle intelligent inspection path planning method and system based on visual inspection

The invention relates to the technical field of unmanned aerial vehicle intelligent inspection path planning, and discloses an unmanned aerial vehicle intelligent inspection path planning method and system based on visual inspection, and the method comprises the steps: collecting an image, and carrying out the differential superposition processing, and obtaining environment visual data; feature correction is extracted to determine a candidate area, and a template is matched to determine the type and severity of an abnormal event; performing local path planning according to an event generation mode switching instruction; according to the trajectory parameters meeting the minimum turning radius constraint, a safe fly-around path is obtained; the method comprises the following steps of: acquiring a real-time image stream, adjusting a speed parameter to obtain an optimized trajectory, processing the real-time image stream and inertial measurement data generated in the optimized trajectory through a visual inertial odometer technology to obtain environment three-dimensional point cloud data, and performing cyclic verification according to the environment three-dimensional point cloud data to obtain an event processing verification result. According to the invention, hierarchical response and dynamic local path re-planning of abnormal events can be realized.
Owner:GUANGDONG CHENGYU ENG CONSULTING SUPERVISION CO LTD

Unmanned aerial vehicle dynamic environment sensing method based on visual large model, electronic equipment and storage medium

The invention relates to an unmanned aerial vehicle dynamic environment sensing method based on a visual large model, electronic equipment and a medium, and the method comprises the steps: carrying out the instance segmentation and optical flow estimation of multiple frames of original images based on the visual large model and an optical flow estimation network, and obtaining a dynamic mask and inter-frame optical flow information; performing feature point classification on the depth information, the dynamic masks and the inter-frame optical flow information corresponding to the multiple frames of images to obtain a classification result; constructing a re-projection error model based on the classification result to obtain an estimated pose; evaluating the quality of the feature points based on the estimated pose to obtain an optimized pose; and constructing a feature point global map based on the world coordinates of the continuous frame road sign points and the optimized poses. Through a multi-information fusion and step-by-step optimization mechanism, the method has relatively high robustness, can adapt to various dynamic environments, and improves environment perception performance and operation reliability in different scenes.
Owner:CIVIL AVIATION UNIV OF CHINA

Techniques for vision-based robot control

Techniques for controlling a robot include receiving sensor data and one or more goal specifications, processing the sensor data, a robot size, and the one or more goal specifications using one or more trained encoders to generate a plurality context tokens, processing the plurality of context tokens using one or more trained decoders to generate a robot plan, and controlling a robot based on the robot plan.
Owner:NVIDIA CORP

Dynamic data tracing method and system for network security

The invention provides a dynamic data traceability method and system oriented to network security, and relates to the technical field of network security and data governance crossing, and the method comprises the steps: carrying out the real-time visual monitoring of a user interaction interface of a data transaction platform, and capturing the visual behavior information related to data calling; extracting feature parameters based on the visual behavior information to obtain a visual feature data set, and performing structured conversion on the visual feature data set to obtain a structured call log record containing visual behavior associated information; based on the structured call log record, extracting a time sequence, a frequency and a spatial distribution feature of a call behavior as basic indexes; a plurality of feature sampling states are dynamically selected in a monitoring time window, and a dynamic feature evaluation set is constructed based on the feature sampling states. According to the invention, the accuracy and controllability of network security protection in a data transaction scene can be improved.
Owner:TAIZHOU DIGITAL GROUP CO LTD

Vision and sensing fusion-based medicine bottle array detection system

The invention provides a medicine bottle array detection system based on vision and sensing fusion, and relates to the technical field of data processing, and the system comprises the steps: obtaining the spatial distribution data of a vision unit and the posture state data of a sensing unit, fusing the position parameters and state parameters of each bottle body, building a multi-dimensional data node under a unified time index, calculating the spatial aggregation degree of the medicine bottle array to obtain an array quantity, combining the aggregation change of each time point into an evolution chain table, carrying out joint operation on the difference result and the state parameter of adjacent nodes in the evolution chain table, calculating the dynamic coupling degree of the medicine bottle array to obtain a sequence potential energy value, and dividing the sequence potential energy values of the different numerical value intervals into a plurality of level nodes, and performing consistency analysis and anomaly detection on features of the different level nodes to obtain detection result data. By detecting the state of the medicine bottles and the arraying direction, bottle toppling early warning is achieved.
Owner:WENZHOU JINGYUE TECH CO LTD

Passable area reasoning method and system based on visual language model

PendingCN121767911AAchieve collaborative understandingEnable high-level semantic reasoningCharacter and pattern recognitionBiological modelsSemantic alignmentVision based
The invention provides a passable area reasoning method and system based on a visual language model, and the method comprises the steps: obtaining the multi-modal data of a vehicle and the current position information of the vehicle; analyzing the multi-modal data, and determining visual features and traffic symbol features; performing spatial position coding on the visual object and the traffic symbol elements, and determining aerial view angle coordinate information; performing semantic alignment on the visual features and the traffic symbol features, and determining a shared embedding representation; constructing a traffic semantic map by fusing, sharing and embedding representation based on a graph neural network and bird's-eye view coordinate information of a visual object and a traffic symbol element; and according to the current position information of the vehicle, the traffic semantic map and a preset traffic rule, generating a bird's-eye view semantic map including a passable area, a no-pass area and a semantic association relationship. According to the method and the device, semantic alignment and consistency expression of visual perception and traffic symbol recognition are realized, and further feasible region reasoning of a complex traffic scene is realized.
Owner:SHANGHAI JIAOTONG UNIV

Injection mold surface defect detection method based on visual inspection

The invention belongs to the technical field of injection mold detection, and discloses an injection mold surface defect detection method based on visual inspection, and the method comprises the following steps: S1, adaptively adjusting collection parameters according to mold process parameters, and ensuring clear collection of different process surface defect features; s2, a deformation matrix is output through the mold thermal deformation finite element model, image registration is guided in combination with a deformation field, and the defect position deviation of the thermal-state mold is corrected; the method comprises the following steps: constructing a vibration-fuzzy kernel mapping model based on vibration sensor data, and restoring a fuzzy image by using an improved Richardson-Lucy algorithm; s3, a process exclusive texture primitive library is constructed, and unified threshold positioning deviation is avoided; s4, converting process parameters into feature extraction weights through a process perception attention CNN model, and accurately capturing defect core features under different processes; according to the design, the detection method can adapt to the surface of a multi-process mold without replacing a model, and the debugging cost of cross-process detection is greatly reduced.
Owner:SUZHOU XINGKAISHENG INTELLIGENT TECHNOLOGY CO LTD

Vision-based cross-network interaction method and system

The invention provides a vision-based cross-network interaction method and system, and relates to the technical field of intelligent interaction, and the method comprises the steps: obtaining a visual interaction sequence, constructing multi-modal feature representation, achieving cross-domain semantic alignment, deconstructing visual information into a hierarchical control instruction set, and transmitting the hierarchical control instruction set to a target network environment for execution after security classification. And a bidirectional mapping relation graph is constructed to realize incremental optimization. According to the method, semantic bridging between heterogeneous networks can be established, the cross-domain control precision is improved, and meanwhile safe interaction in a network isolation environment is guaranteed.
Owner:ZHONGTIAN ZHILING (BEIJING) TECH CO LTD

Social Networking Content Supplemented Web Page Linker

A modular system designed for privacy-preserving content recognition and supplemental content delivery across web and mobile environments. The system employs lightweight character sampling and vision-based recognition to generate unique content fingerprints without storing or replicating original data. It features a hybrid processing architecture, using local computing resources for intensive tasks while optimizing performance on resource-constrained devices. Core functionalities include multi-method content fingerprinting, real-time monitoring with adaptive sampling, and secure supplemental content association. Operating entirely on the client-side, it complies with website terms of service and privacy regulations. Advanced features include AI-driven content recognition, blockchain-based verification, and granular content targeting through resizable selection interfaces. This technology enables seamless delivery of supplemental content while preserving privacy, reducing resource usage, and ensuring scalability across browsers, mobile applications, and edge devices. It is particularly applicable in industries such as education, retail, and secure data sharing.
Owner:TORRES TERRY LEE