Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

81812 results about "Computer graphics (images)" patented technology

Computer graphics are pictures and films created using computers. Usually, the term refers to computer-generated image data created with the help of specialized graphical hardware and software. It is a vast and recently developed area of computer science. The phrase was coined in 1960, by computer graphics researchers Verne Hudson and William Fetter of Boeing. It is often abbreviated as CG, though sometimes erroneously referred to as computer-generated imagery (CGI).

Image processing apparatus and method, and storage medium

A pattern which does not appear at a flat portion in normal binarization processing is set as a code pattern, and a code formed from this pattern is attached. At this time, code attachment with little degradation in image quality is implemented by selecting an unnoticeable pattern.
Owner:CANON KK

Three-dimensional environment reconstruction optimization method based on multi-sensor fusion data

The invention discloses a three-dimensional environment reconstruction optimization method based on multi-sensor fusion data, and relates to the field of three-dimensional environment reconstruction optimization, and the three-dimensional environment reconstruction optimization method based on the multi-sensor fusion data comprises the following steps: S1, collecting multi-source sensor data, and constructing a data set under a unified coordinate system; s2, generating dense visual point cloud, and extracting laser point cloud features to construct a model; s3, establishing a local three-dimensional model, and generating a local environment image; s4, shadow parameters are extracted through shadow geometric analysis, and time sequence optimization is carried out; s5, consistency verification and correction are carried out, and three-dimensional reconstruction data are output; and S6, comparing the reconstruction data with the navigation map database, and carrying out map optimization updating. According to the method, time synchronization and space calibration are carried out on data acquired by the depth camera and the laser radar, complete and accurate three-dimensional information modeling of the target environment is realized, and the geometric precision of environment reconstruction and the image detail reduction capability are improved.
Owner:NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER

Method for automatically drawing OpenGL program by using Vulkan

The invention discloses a method for automatically drawing an OpenGL (Open Graphics Library) program by using Vulkan. The method comprises the following steps of: creating a context used by the Vulkan, initializing each module, processing an OpenGL instruction related to texture and data buffering, and managing storage of texture and data buffering resources in a video memory; a shader program used by the OpenGL is preprocessed into a format acceptable to Vulkan, and an OpenGL shader program instruction is created and destroyed; processing an OpenGL (Open Graphics Library) instruction related to frame buffering to generate structural body information required by Vulkan dynamic rendering; an OpenGL instruction of the sampler is also created; processing an OpenGL (Open Graphics Library) instruction for creating a vertex input format and managing a vertex data buffer area, and maintaining vertex input information, a vertex buffer area and an index buffer area required by Vulkan; and finally, drawing or calculating, distributing and calling Vulkan on the basis of all the instructions.
Owner:ZHEJIANG UNIV +1

Holder tracking method and device based on binocular camera, and storage medium

The invention discloses a cradle head tracking method and device based on a binocular camera and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: processing image data based on a binocular parallax principle, and generating a three-dimensional coordinate of a center point of a tracking target; based on the three-dimensional coordinates of the camera coordinate system and the offset from the optical center of the camera to the rotation center of the holder, generating three-dimensional holder coordinates through coordinate transformation solution; determining a historical track based on the tracking target feature information and a historical target feature matching result, and outputting an identifier and a three-dimensional position observation value through correlation verification of a three-dimensional holder coordinate and the historical track; inputting an observation updating equation correction state through the identifier and the three-dimensional position observation value, and outputting a three-dimensional prediction position; based on the three-dimensional prediction position and a deviation formula, calculating the angle deviation with the camera image center under the holder coordinate system, and driving the holder to center the target in the picture center according to the angle deviation. The problem that the target tracking effect is poor is solved, and the robustness of target tracking in a complex scene is improved.
Owner:SHENZHEN EMEET TECH CO LTD

Three-dimensional reconstructions based on gaussian primitives

In implementation of techniques for three-dimensional reconstructions based on Gaussian primitives, a computing device implements a reconstruction system to receive a first digital image depicting an object from a first angle and a second digital image depicting the object from a second angle. The reconstruction system segments the first digital image and the second digital image into patches. The reconstruction system then generates, using a machine learning model, three-dimensional Gaussian primitives that predict parameters of points of the object in a three-dimensional space that correspond on a per-pixel basis to pixels of the patches. The reconstruction system then forms a three-dimensional reconstruction of the object for display in a user interface by merging the three-dimensional Gaussian primitives.
Owner:ADOBE INC

Three-dimensional dynamic scene reconstruction method and apparatus, and storage medium

The present disclosure relates to the field of computer vision and discloses a three-dimensional dynamic scene reconstruction method and apparatus, and a storage medium. The three-dimensional dynamic scene reconstruction method comprises: acquiring synchronized videos of a plurality of viewpoints of a dynamic scene; computing matching points between video images of different viewpoints, and estimating intrinsic and extrinsic parameters of each camera; obtaining a Gaussian splatting point set {p0} on the basis of a sparse point cloud constructed according to the depth of each matching point; for the first image frame of each video, using {p0} to perform static training thereon, to obtain a Gaussian splatting point set {p}; for the remaining image frames, dividing {p} into a static point set {S} and a dynamic point set {D}, performing dynamic training on {D}, and constructing a dynamic Gaussian splatting point set {P} from {p}, {S}, and the final {D}; and, in view of the intrinsic and extrinsic parameters of each camera, rendering {P} using a Gaussian splatting rendering pipeline, to obtain rendered images at different moments from new viewpoints.
Owner:TSINGHUA UNIVERSITY

Intelligent shooting method for scene understanding and script analysis driven by large science and technology movie and television model

The invention relates to an intelligent shooting method for scene understanding and script analysis driven by a large science and technology movie and television model, and belongs to the technical field of machine vision. The method comprises the steps that script text features are extracted and decomposed to obtain plot development, emotional fluctuation and artistic style information, a reference film and television work with the maximum overall matching score is selected based on a decomposition result, and an optimal shooting strategy vector is extracted; analyzing the shot scene to generate visual features; fusing the text features and the visual features to obtain a multi-modal semantic representation, and generating a dynamic scene-task knowledge graph according to the optimized multi-modal semantic representation so as to generate a shot scheduling strategy; the shooting process is tracked, the shooting sequence and the lens application mode are monitored in real time, and when it is detected that the shooting sequence deviates, lens connection deviates or visual expression does not conform to expectation, an intelligent optimization mechanism is triggered; and when the deviation exceeds a set threshold value, a manual intervention prompt is given out. The shooting cost can be reduced, and the manufacturing efficiency can be improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

System and method for reconstructing 3D scene data from 2D image data

A method and apparatus for reconstructing a three-dimensional (3D) scene from a two-dimensional (2D) input image of the scene using a fully-differentiable transformer-based encoder-decode. A 2D input image encoded into a set of image features using a pre-trained vision transformer model, wherein the vision transformer model is pre-trained with multi-view RGB image supervision and point cloud supervision. The set of image features is projected onto a 3D triplane representation using a transformer decoder to obtain output triplane tokens. A triplane representation is created from the tokens and queried. 3D point features of color and density for volumetric rendering re predicted using a multi-layer perceptron. The geometry of the generated 3D asset is represented with a surface mesh including vertices and triangular faces. A texture map by is created with a multichannel image in UV space. Multiple views of the 3D scene are simultaneously generated based on the surface mesh.
Owner:FUTUREVERSE IP LTD

Tooth three-dimensional modeling system based on computer vision, computer equipment and readable storage medium

The invention relates to the technical field of tooth modeling, and discloses a three-dimensional tooth modeling system based on computer vision, computer equipment and a readable storage medium. According to the method, mirror reflection, diffuse reflection and subsurface scattering components in an original image are separated, mirror reflection intensity is normalized in combination with a dynamic truncation algorithm, pixel saturation is eliminated, groove and nest textures are reserved, a complete point cloud is obtained based on a two-dimensional texture image and cubic spline repair, and a multi-exposure point cloud sequence is obtained through bimodal calibration. The method comprises the following steps: solving the problem of data dislocation, carrying out weight assignment and data fusion on three-dimensional points in a plurality of exposure point cloud sequences to obtain three-dimensional fusion feature data, combining layered optical modeling and photon tracking compensation deviation, fusing clinical constraints, finally dynamically adjusting parameters, feeding back and optimizing, and outputting a micron-sized precision model. The modeling defect caused by difficulty in effectively coordinating feature contribution degrees under different exposure conditions is overcome, and high-precision modeling is realized.
Owner:SHENZHEN JINSHI LIMEI MEDICAL TECH CO LTD

Point cloud segmentation method and apparatus, storage medium, and electronic device

The present application discloses a point cloud segmentation method and apparatus, a storage medium, and an electronic device. The method comprises: for a target sample point in a spatial point cloud, determining a curvature feature of a local plane corresponding to the target sample point in the spatial point cloud; on the basis of the curvature feature, determining a local density corresponding to the local plane; on the basis of candidate relative distance values between a plurality of other sample points and the target sample point, determining a target relative distance value corresponding to the target sample point, wherein the plurality of other sample points are sample points other than the target sample point in the spatial point cloud; and on the basis of the local density and the target relative distance value, segmenting the spatial point cloud to obtain a spatial point cloud segmentation result.
Owner:CHINA TELECOM BESTPAY CO LTD

Methods for adjusting and / or controlling immersion associated with user interfaces

In some embodiments, an electronic device emphasizes and / or deemphasizes user interfaces based on the gaze of a user. In some embodiments, an electronic device defines levels of immersion for different user interfaces independently of one another. In some embodiments, an electronic device resumes display of a user interface at a previously-displayed level of immersion after (e.g., temporarily) reducing the level of immersion associated with the user interface. In some embodiments, an electronic device allows objects, people, and / or portions of an environment to be visible through a user interface displayed by the electronic device. In some embodiments, an electronic device reduces the level of immersion associated with a user interface based on characteristics of the electronic device and / or physical environment of the electronic device.
Owner:APPLE INC

System and method of three-dimensional object cleanup and text annotation

Some examples of the disclosure are directed to object manipulators and associated processes for manipulating an object representation in a three-dimensional environment. The object representation may correspond to a scan of a real-world object in a real-world environment. The object manipulators may include an object cleanup manipulator and a text annotation manipulator. The object cleanup manipulator may be selectable to display one or more control affordances providing functionality for selectively removing portions of the object representation in the three-dimensional environment and / or selectively adjusting one or more parameters of the object representation in the three-dimensional environment. The text annotation manipulator may be selectable to display one or more control affordances providing functionality for selectively generating one or more text labels in the three-dimensional environment. The one or more text labels may be associated with the object representation in the three-dimensional environment.
Owner:APPLE INC

Urban building three-dimensional automatic modeling and visualization method

The invention discloses an urban building three-dimensional automatic modeling and visualization method, and belongs to the technical field of building three-dimensional modeling. The method comprises the steps that point cloud data, high-resolution images and geographic information system data of urban buildings are acquired, data cleaning, registration and alignment are carried out, and preliminary building digital representation is formed; accurately segmenting each building, and identifying the contour and main structural features of the building; based on the data integrity and the building complexity, adaptively selecting a proper reconstruction strategy to carry out three-dimensional reconstruction; in the reconstruction process, the geometric structure is analyzed and optimized in real time, and potential topological problems are repaired; automatically generating missing details based on a predefined architectural style library and a component library, and performing material inference and texture mapping; a graph structure is used for representing the relation between the buildings, and the positions and orientations of the buildings are adjusted through a global optimization algorithm; a rendering engine supporting multi-level detail switching is developed, and smooth visualization and interaction of a large-scale city scene are achieved.
Owner:CHANGZHOU JINTAN DISTRICT LUOSUI TECHNOLOGY CO LTD

Intelligent event identification method and system based on high-speed camera

The invention provides an intelligent event identification method and system based on a high-speed camera, and the method comprises the steps: setting the collection parameters of the high-speed camera, and triggering the camera to collect a target scene video stream. And hardware acceleration decoding processing is carried out on the collected original video data stream, real-time environment illumination information of the environment illumination sensor is obtained, and dynamic brightness equalization processing is executed. And performing motion adaptive denoising processing on the video sequence. Geometric distortion correction is carried out on the image sequence through camera calibration parameters, sub-pixel-level displacement vectors and dense optical flow field data of a moving object are extracted, and multi-scale morphological features are extracted. And the features are fused to generate motion feature data, the data are processed through a spatio-temporal joint event classification model, an event identification result is output, the result is bound with a high-precision timestamp, and event identification information is output to an industrial control system display device in real time. According to the invention, the accuracy and real-time performance of event identification can be improved.
Owner:广州思林杰科技股份有限公司

Quantification method for microscopic cracks inside 3D printed concrete and system thereof

Provided are a quantification method for microscopic cracks inside 3D printed concrete and a system thereof. Before measuring the microscopic crack images according to relevant standards, firstly, a denoising network, improved attention-guided denoising neural network (IADNet) is adopted. IADNet can extract features from microscopic crack images from different perspectives, perceive noise from multiple levels, and perform denoising processing, greatly improving image quality and enhancing texture details, which is beneficial for training segmentation networks. The combination of IADNet and semantic segmentation algorithm has the ability to finely recognize image information, quantify microscopic crack recognition, overcome the shortcomings of measurement and analysis of microscopic cracks inside 3D printed concrete, and improve construction efficiency and quality.
Owner:JIANGXI COMMUNICATIONS INVESTMENT GROUP CO LTD +2

Systems and Methods for Latent Hyperspace Navigation in Spatiotemporal Media

A system and method for latent hyperspace navigation in spatiotemporal media using hierarchical and Lorentzian autoencoders. The system compresses spatiotemporal media into navigable latent representations while preserving geometric and semantic relationships through tensor structure maintenance. A latent hyperspace manager organizes compressed representations as geodesic trajectories within a geometric manifold structure based on differential geometry principles. A geodesic trajectory mapper computes optimal navigation paths through the high-dimensional space, while symbolic anchors positioned at semantically significant locations serve as persistent reference points. Spatiotemporal routing protocols manage navigation decisions across multiple temporal scales. A strategy caching system preserves successful navigation patterns for reuse, enabling continuous learning. The system generates synthetic content during navigation to support infinite zoom capability, allowing exploration beyond original media boundaries. Cross-modal fusion combines diverse input modalities into unified representations, applicable to immersive media exploration, scientific visualization, and surveillance analysis.
Owner:ATOMBEAM TECH INC

Image restoration and super-resolution reconstruction system and method based on deep learning

The invention provides an image restoration and super-resolution reconstruction system and method based on deep learning, and belongs to the technical field of digital image processing. The invention aims to solve the problems of high calculation complexity and resource consumption, limitation of long sequence processing, high training difficulty and texture scene deficiency when a multi-scale residual network based on a Transform architecture is used for image resolution conversion. The reconstruction system comprises: an image preprocessing module performing window division and video memory optimization on an input low-resolution image; the multi-layer fusion network dynamically adjusts the characteristics of the low-resolution image, captures channel information in different scenes, performs interactive fusion, performs comparison supervision, establishes an information communication channel, dynamically adjusts and optimizes parameters through negative feedback, and obtains a super-resolution image. And the loss function module maximizes the similarity of the super-resolution image and the high-resolution image in the segmentation feature space to obtain a final super-resolution image.
Owner:QIQIHAR UNIVERSITY

Semantic segmentation method for low-resolution road scene

The invention discloses a semantic segmentation method for a low-resolution road scene, and aims to solve the problems of difficulty in small target recognition, fuzzy details, texture information loss and the like existing in a low-resolution image in the conventional semantic segmentation technology. The method comprises the following steps: (1) collecting a low-resolution road scene image and a corresponding semantic tag; (2) constructing a semantic segmentation model consisting of an edge guidance module (BGM), a double-domain feature decomposer (DDFD), a domain alignment attention fusion module (DAAFM) and a double-layer attention context aggregation module (HACAM); (3) designing a joint loss function to carry out multi-scale supervision on semantic regions, edges and middle features; (4) carrying out model training by utilizing the road scene image; and (5) outputting a semantic segmentation result map and an edge prediction map. The boundary perception capability is enhanced by introducing learnable pixel difference convolution, the extraction precision of a small target and a global structure is improved by combining frequency domain and spatial domain feature alignment, and context semantic relationship expression is optimized by fusing a channel and a spatial attention mechanism. The method effectively improves the semantic segmentation precision and boundary restoration capability of the model in a low-resolution complex road environment, and is suitable for intelligent analysis tasks of road images in scenes of automatic driving, intelligent traffic, severe weather and the like.
Owner:CENT SOUTH UNIV

System and method for extracting three-dimensional gluing contour of shoe sole based on visual single-line laser

The invention relates to the technical field of computer vision and industrial automation, in particular to a shoe sole three-dimensional gluing contour extraction system and method based on vision single-line laser, and aims to solve the problems that virtual calibration target spots cannot be accurately generated based on shoe sole geometry, the positions and sizes of the target spots are difficult to determine by combining curvature extreme values and principal component analysis in the prior art, and the production cost is low. The problem that a double-branch deep learning model cannot be adopted to fuse feature prediction transformation, and the re-projection error is increased is solved; a virtual calibration target spot is automatically generated based on sole geometry through a feature fusion calibration module, a grid is generated through point cloud processing and Poisson reconstruction, the position and size of the target spot are determined by combining a curvature extreme value and principal component analysis, a corresponding relation is established by utilizing two-dimensional and three-dimensional feature matching, initial alignment is realized through ICP and re-projection error optimization, and the target spot position and size are determined. A double-branch deep learning model is adopted to be fused with feature prediction transformation, iterative optimization is carried out through space consistency errors, and re-projection errors are reduced.
Owner:ANHUI UNIV

Non-standard intelligent customized welding system based on vision and surface gradient

The invention relates to the field of welding, and discloses a non-standard intelligent customized welding system based on vision and surface gradient, which comprises a visual perception module used for iteratively approaching a workpiece from a preset initial height through a self-adaptive high angle shooting exploration mechanism, and combining layered grid division based on a camera view and a progressive multi-angle expansion scanning mode, collecting multi-view three-dimensional point cloud data of the workpiece; and the point cloud processing module is used for performing regional progressive registration on the multi-view three-dimensional point cloud data, implementing threshold constraint based on point cloud curvature characteristics by dynamically adjusting registration step length, and eliminating accumulative errors in combination with pose map optimization. Through self-adaptive high-angle shooting and layered grid scanning, multi-view point cloud can be automatically collected without manually presetting a workpiece model, regional registration and curvature constraint are combined, the workpiece model is constructed, a weld joint structure is automatically extracted based on surface gradient features, traditional manual labeling is replaced, a welding track is generated according to weld joint topology, and the precision is dynamically corrected.
Owner:SHANGHAI SHENGSHI WEISHENG TECH CO LTD

Flexible display module surface defect image recognition method

The invention relates to the technical field of industrial product surface quality detection, in particular to a flexible display module surface defect image recognition method, which comprises the following steps: acquiring a plurality of surface images of a flexible display module under different light sources and carrying out distortion removal processing on the surface images; reconstructing and generating three-dimensional reference point cloud data representing the current curved surface form of the module; re-projecting the distorted image to the ideal rigid plane according to the relationship, generating a plurality of corrected images, and generating a plurality of corrected images to eliminate geometric and luminosity distortion introduced by flexible deformation; obtaining a defect candidate area binary image; and extracting a multi-dimensional feature vector of the defect candidate region from the binary image of the defect candidate region, and classifying the feature vector by using a pre-trained defect classification model to obtain a defect identification result. Through the three-dimensional reference point cloud reconstruction and image re-projection technology, the problem of geometric distortion caused by surface deformation of the flexible display module is effectively solved, and misjudgment and missed judgment are avoided.
Owner:HUNAN HUICHENGXIN TECHNOLOGY CO LTD

Pet target detection method and device and camera

The invention relates to the technical field of target detection, and discloses a pet target detection method and device and a camera, and the method comprises the steps: carrying out the motion triggering collection and image enhancement preprocessing of a front end region of a feeder, and obtaining an enhanced image frame sequence; performing feature extraction of dynamic receptive field adjustment on the enhanced image frame sequence to obtain pet feature descriptors and position information; behavior time sequence feature analysis is executed, and pet behavior sequence feature vectors are obtained; constructing a state transition diagram according to the pet behavior sequence feature vector, and performing time sequence consistency analysis to obtain a pet state judgment result; power management and decision execution are carried out on the feeder based on the pet state judgment result, feeding control under the low-power-consumption condition is achieved, behavior misjudgment caused by posture fluctuation is effectively avoided, the behavior recognition accuracy is improved, the accurate feeding control problem in a multi-pet family is solved, and the user experience is improved.
Owner:SHENZHEN ANKED SHITONG ELECTRONICS CO LTD

Visual Transform-based dynamic screening medical image target tracking method and device

The invention provides a dynamic screening medical image target tracking method and device based on visual Transform, and relates to the technical field of computer vision, and the method comprises the steps: standardizing near-infrared or visible light fundus video frames into uniform resolution, constructing a template-search frame pair, and then jointly mapping the two frames of images into a Token sequence; a dynamic local interaction module is embedded in front of each pruning layer of the whole network, a local context is captured by using depth separable convolution and point convolution, a dynamic convolution kernel generator is driven, and a neighborhood Token is adaptively weighted and aggregated. Next, the Token screening and the compression mechanism TSC are operated in the same pruning layer, only the Top-K key Token is reserved, the redundant Token is cut off, and the original index is recorded; the objective of the invention is to improve the positioning stability and reasoning efficiency of a focus area (such as an optic disc) in a complex operation video.
Owner:XIAMEN UNIV OF TECH

Building engineering progress automatic identification and early warning system based on computer vision

The invention discloses a building engineering progress automatic identification early warning system based on computer vision, which relates to the field of building engineering informatization and comprises a synchronous calibration module, a joint calibration module, a pose generation module, a mapping construction module, a resampling module, a mapping registration module and a comparison early warning module. Clock synchronization and rolling readout calibration are carried out on a camera and an inertial measurement unit, a continuous time pose is established in a frame, sampling is carried out according to rows, and an equivalent global shutter frame is generated in combination with plane and pixel-level geometric mapping; outputting a camera pose track and a local measurement map by adopting back-end optimization containing a rolling mechanism, and registering with the building information model; component detection, instance segmentation and state discrimination are completed on an equivalent global shutter frame, pixel domain measurement is unified into engineering quantity under pose and registration constraints, the engineering quantity is mapped to a work decomposition structure and a progress plan, deviation is calculated, and graded early warning is output according to a multi-threshold rule. Long-term stable and reliable operation and evidence traceability of the system are guaranteed through whole-course quality monitoring and threshold grading.
Owner:JIANGXI TRANSPORT CONSULTATION +1

Object monitoring method based on multi-camera joint calibration technology and monitoring camera thereof

The invention relates to the technical field of vision, in particular to an object monitoring method based on a multi-camera joint calibration technology and a monitoring camera thereof. The method comprises the following steps: driving all cameras to synchronously shoot a calibration reference object with known geometric characteristics, and solving internal and external parameters of each camera and an accurate space pose relationship between the internal and external parameters by utilizing shot images to form a joint space relationship chain; controlling the camera array to synchronously shoot a target object, and correspondingly converting the two-dimensional image points acquired by the cameras into spatial point coordinates in a unified three-dimensional coordinate system by using the joint spatial relation chain to generate three-dimensional point cloud data of the target object; and processing the generated three-dimensional point cloud data, extracting key geometric features of the target object in real time, performing comparative analysis on the extracted features and a preset standard or a historical state, and outputting a state monitoring result of the target object. According to the invention, the influence of temperature drift and mechanical vibration on the measurement precision is effectively suppressed, and the stability of long-term monitoring of an industrial field is guaranteed.
Owner:SUZHOU MEILITO ELECTRONIC TECH CO LTD

Style transfer using generative diffusion features

The present invention sets forth techniques for performing style transfer from multiple supplied style images to a supplied content image to generate novel images that include style elements from the multiple supplied style images and content elements from the supplied content image. The techniques include guiding one or more self-attention and cross-attention layers included in a machine learning model based on the multiple supplied style images, such that content elements and style elements included in the style images are not entangled when generating the novel images. The techniques also distill a small subset of representative attention map values from multiple style images, improving performance while reducing computational costs compared to processing all attention map values from the multiple style images.
Owner:DISNEY ENTERPRISES INC

Welded pipe surface defect detection method based on robot visual inspection

The invention relates to the field of image processing, and particularly discloses a welded pipe surface defect detection method based on robot visual inspection. The method comprises the steps that a robot carries a binocular camera and an annular LED light source and moves at a constant speed in the axial direction of a welded pipe to collect orthographic and inclined views, and a three-dimensional point cloud is constructed; and establishing a parameterized mapping function based on the point cloud, and converting the 3D coordinate into a 2D expansion surface coordinate. In the convolutional neural network, a first layer is inserted into a spatial transformation network to correct distortion of the expanded image, deformable convolution is adopted to extract edge, local deformation and specific defect response features, and standard convolution is combined to extract global features; and fusing multi-scale features and adding an attention mechanism to improve the weight of a defect region, and outputting a defect category and a bounding box offset after generating a candidate box. The method effectively solves the problems of stretching, deformation and defect distortion of welded pipe curved surface imaging, reduces the imaging difference of the same defect, and remarkably improves the defect positioning precision and recognition accuracy.
Owner:JINAN HENGPENG MACHINERY CO LTD

Lightweight multi-modal content identification system based on double-track migration framework

The invention discloses a lightweight multi-modal content recognition system based on a double-track migration framework, and relates to the technical field of content recognition, and the system comprises a data collection module which is used for synchronously collecting multi-source content of a text and an image and carrying out standardization processing and tensor construction to form a fusion tensor X; and the model construction module is used for inputting the fusion tensor into a dual-track migration structure constructed based on a Transform backbone network, and the dual-track migration structure realizes task semantic alignment and structure migration under parameter freezing through Prompt Learning embedding and Adapter-Tuning insertion, and outputs an intermediate representation of modal alignment. According to the method, the training and deployment cost of multi-modal content recognition is remarkably reduced, the semantic expression ability in the modal is enhanced, the cross-modal alignment precision and fusion depth are effectively improved, and the perception and recognition ability of the model to the complex semantic relationship is enhanced.
Owner:CCTV INT NETWORK CO LTD

Marker writing system

Embodiments of the invention provide a pen-stylus that comprises a first antenna having a hemispherical shape at a proximal end of the pen-stylus where the pen-stylus engages with a display of a tablet device. The pen-stylus also includes a second antenna, wherein the first antenna and the second antenna communicate with the tablet device to enable a pen-stylus user to draw on the display of the tablet device. The pen-stylus additionally includes an insulating material between the first antenna and the second antenna that prevents interference between the first antenna and the second antenna. In some embodiments, the first antenna and the insulating material reside in a marker tip attached to the pen-stylus, the marker tip transmits axial forces received by engagement of the pen-stylus with the display on the tablet to a force sensor. The pen-stylus further includes a force sensor that converts axial forces received from the marker tip into an electrical signal and transmits the electrical signal to the tablet by at least one of the first antenna and the second antenna.
Owner:REMARKABLE AS