Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

134 results about "Visual saliency" patented technology

Visual salience (or visual saliency) is the distinct subjective perceptual quality which makes some items in the world stand out from their neighbors and immediately grab our attention.

Two-way visual saliency detection method and device combining difference guidance and texture enhancement

The invention discloses a two-way visual saliency detection method and device combining difference guidance and texture enhancement, and the method comprises the following steps: 1, obtaining an image to be subjected to saliency detection, and carrying out the marking and preprocessing of a data set; 2, constructing a visual saliency detection model which comprises a dual-path encoder (a saliency detection path and an image reconstruction path), an adaptive interaction network, a decoder network and an output network; a significance detection path in the dual-path encoder network uses a pre-trained ConvNeXt encoder, and an image reconstruction path uses VQ-VAE as a backbone network; the adaptive interactive network comprises a multi-scale convolution module and a gating fusion module; the decoder network comprises a mutual conversion attention module and a double-gating fusion module; the output network comprises a multi-level feature fusion module; 3, training the saliency detection model to obtain a trained saliency detection model; and 4, carrying out saliency detection on the image data by adopting the trained saliency detection model.
Owner:SICHUAN UNIV

Short video intelligent editing method and system based on multi-modal analysis

The invention discloses a short video intelligent editing method and system based on multi-modal analysis, and relates to the technical field of video editing. The method is used for improving editing efficiency and visual experience and comprises the following steps: extracting lip motion features of a character, visual saliency features of a commodity and a voice emotion intensity value from a target short video stream to form multi-modal time sequence data; afterwards, the voice stream is recorded, a product keyword timestamp is extracted, the alignment degree is calculated through dynamic time warping in combination with a visual saliency peak value, and a preliminary editing point set is generated through weighted evaluation in combination with an emotional intensity value; constructing an editing decision optimization model based on deep reinforcement learning, taking the multi-modal features as state input, adjusting the retention probability of editing points through a joint reward function, and selecting an optimal transition mode; and the lip movement and voice synchronization error before and after the editing point and the emotional and visual continuity of the transition section are analyzed, the discontinuous region is smoothed, and the edited finished product is output, so that precise short video intelligent editing is realized.
Owner:ANHUI XINGBANG DIGITAL TECHNOLOGY GROUP CO LTD

Real-time rendering and interaction method for immersive virtual reality scene

The invention relates to the technical field of computers, and discloses a real-time rendering and interaction method and system for an immersive virtual reality scene. The method comprises the following steps: fusing tuner inertial data and eyeball tracking data, and constructing a prospective state prediction model; generating a predictive focus field in combination with scene visual saliency; synthesizing an anisotropic temporal-spatial resolution graph according to the predicted head angular velocity; gPU variable-rate coloring is driven to realize non-uniform rendering; and re-projection or dynamic fuzzy correction is executed in a self-adaptive manner according to the attitude prediction error before display. According to the technical scheme, the perception delay and the rendering load are remarkably reduced, and the frame rate stability and the visual immersion in a high-dynamic scene are improved.
Owner:CHENGDU TECHNICIAN COLLEGE (CHENGDU VOCATIONAL & TECH COLLEGE OF IND & TRADE CHENGDU ADVANCED TECH SCHOOL CHENGDU RAILWAY ENG SCHOOL)

Remote sensing image text retrieval method based on remote sensing multi-modal basic model

The invention relates to the technical field of remote sensing image analysis and cross-modal retrieval. The invention discloses a remote sensing image text retrieval method based on a remote sensing multi-modal basic model, which applies the large-scale pre-training capability of a CLIP large model to semantic alignment of a remote sensing image and a text by finely adjusting the CLIP large model. By introducing the visual saliency calculation module and the visual block fine-grained selection integration module, the problems of multi-scale targets and redundant information in the remote sensing image are effectively solved, fine-grained semantic alignment between the image and the text is realized, and the retrieval accuracy is improved. Particularly, under the condition that the image contains a plurality of salient targets and redundant regions, the cross-modal semantic alignment fine-grained filtering method provided by the invention can accurately identify key information blocks in the image and perform fine matching with text description.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Eye movement tracking method and system based on multi-modal fusion

The invention discloses an eye movement tracking method and system based on multi-modal fusion, and particularly relates to the technical field of eye movement tracking, and the method comprises the steps: collecting a face image, an eye image and a head movement parameter of a user through an RGB camera, an infrared camera and an IMU which are synchronously arranged; the signals are processed in parallel, sight line direction vectors based on the face and the eyes and reliability measurement of the sight line direction vectors are extracted respectively, and the head space posture is calculated; dynamically selecting high-confidence cooperation, master-slave compensation or conflict arbitration strategies according to reliability measurement, and fusing to generate an initial sight direction; and mapping the direction to a display plane to obtain an initial fixation coordinate, performing optimization calibration in combination with visual saliency analysis of screen content, and outputting a final fixation point coordinate. The system comprises a signal acquisition and processing module, a parallel computing module, a multi-strategy fusion decision-making module and an output optimization module. According to the invention, through multi-modal information fusion and dynamic strategy selection, the precision and robustness of eye movement tracking are significantly improved.
Owner:南通诺瞳奕目医疗科技有限公司 +1

DR image automatic enhancement method based on multi-scale fusion

The invention discloses a DR image automatic enhancement method based on multi-scale fusion, and the method comprises the steps: carrying out the guided filtering operation of an original DR image through an edge-preserving guided filter, and carrying out the frequency domain layering of different detail information in the original DR image; then, a visual saliency modeling method based on frequency domain residual errors is adopted to evaluate the structural importance of each frequency domain layer, and self-adaptive weighted fusion of multi-scale structural information is achieved; and finally, aiming at detail visibility difference caused by brightness level change, introducing a brightness mapping strategy to generate a multi-channel brightness image set, respectively performing frequency domain enhancement processing, and through collaborative fusion of a spatial domain scale and a frequency domain scale, improving detail discernibility of a weak structure region in an original DR image, and improving the accuracy of the DR image. And automatic enhancement and adaptive optimization of the original DR image under different gray scale structures are realized.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA +1

Eye movement distortion detection method

The invention relates to the technical field of augmented reality glasses optical correction, and specifically provides an eye movement distortion detection method. The method comprises the following steps: acquiring an AR glasses playing image and identifying a visual saliency region to establish an attention priority order; detecting an eyeball movement track direction vector and an instantaneous speed in real time; determining a predicted attention target according to the motion direction and the attention priority; generating a distortion correction parameter set in combination with eyeball arrival time and spatial distribution information; performing pre-distortion compensation rendering on the candidate area in a pre-generation period before eyeballs reach the actual attention area to form a pre-corrected image buffer area; and finally, outputting a pre-corrected image matched with the actual stay area. Space-time synchronization of eyeball movement and image compensation is achieved through predictive rendering, the image distortion and color separation phenomena in the glancing process are eliminated while the low-delay characteristic is maintained, and the visual definition and comfort in a dynamic scene are remarkably improved.
Owner:SHANGHAI JINGYUE LANTU OPTOELECTRONICS CO LTD

Ultra-high-definition video stream adaptive coding method based on deep learning visual saliency

The invention discloses an ultra-high-definition video stream adaptive coding method based on deep learning visual saliency, and the method comprises the steps: carrying out the five-scale Gaussian filtering processing and image pyramid construction of a video frame, and combining Sobel gradient, Laplacian edge and local binary pattern feature extraction to generate a multi-scale feature map; a pre-training saliency detection network is adopted, and a smooth saliency thermodynamic diagram is generated through processing of a feature adaptation layer, a residual encoder, a self-attention mechanism and a transposed convolution decoder; dividing the video frame into a high region, a middle region and a low region according to the saliency thermodynamic diagram, and establishing a regionalization coding parameter table; performing differentiated prediction modes, motion estimation and quantization strategies on different salient regions; and organizing coded data according to an H.265 / HEVC standard, and embedding the saliency thermodynamic diagram into supplementary enhancement information for transmission. According to the method, the important region concerned by the user can be intelligently identified, a differentiated coding strategy based on content semantics is realized, and the coding efficiency is remarkably improved.
Owner:CHANGSHA CHAOCHUANG ELECTRONICS TECH

Ship bollard identification method and system based on adaptive multi-scale multi-grid division

The invention discloses a ship bollard identification method and system based on adaptive multi-scale multi-grid division, and belongs to the technical field of computer vision and image processing. The system divides an image foreground region and a background region through a visual saliency calculation module, carries out dense sampling in the foreground region and sparse sampling in the background region by adopting a non-uniform grid generation module, screens candidate grids in combination with gradient direction consistency and texture features, fuses candidate frames of different scales through a multi-scale image pyramid, and finally obtains a multi-scale image. And accurate positioning of the bollard is realized through the fine grid accurate positioning module. The method solves the problem that the prior art is insufficient in adaptability to ship size, shooting distance and resolution change, improves the precision and generalization ability of bollard positioning in a complex scene, and is suitable for bollard detection scenes of various ship images.
Owner:昆山市交通运输综合行政执法大队 +1

Coding method and system for on-site video return

The invention provides an on-site video return coding method and system, and the method comprises the steps: collecting a video stream of an on-site scene, judging whether a current video frame in the video stream responds to a scene switching video frame or not, and marking the current video frame as a key frame when the current video frame responds to the scene switching video frame; after the current video frame is identified as a key frame, visual saliency analysis is performed on the key frame, and an entropy weight matrix representing regional information importance distribution in the key frame is constructed according to a visual saliency analysis result and texture information entropies of different blocks in the key frame; adjusting code rate allocation weights of different blocks in the key frame according to the entropy weight matrix and a code rate regulation and control strategy of the key frame to obtain adaptive coding configuration adaptive to content characteristics of the key frame; and performing optimization coding on the key frame through adaptive coding configuration to obtain a target code stream in response to a field video return demand. By adopting the scheme of the invention, the dynamic differential coding of the key video frame in a complex scene can be realized.
Owner:SHENHUA RAIL & FREIGHT WAGONS TRANSPORT

Evaluation method and device for model generation image and storage medium

The invention discloses a model generation image evaluation method and device and a storage medium, and relates to the technical field of image processing. According to the method, the image generated based on the text cue word is obtained, wherein the image comprises the image element corresponding to the text cue word; filtering the image to obtain a filtering response value of each image pixel in the image, and generating a visual attention map of the image based on the filtering response values, the visual attention map representing frequency domain energy distribution of the image; and determining a visual saliency value of the image element according to the visual attention map, and determining a layout score of the image according to the visual saliency value of the image element. According to the method, spatial relations such as distances, alignment modes and hierarchical structures among elements are quantitatively evaluated through visual saliency values, and the defect that a previous model cannot accurately evaluate and adjust image layout is overcome.
Owner:SHENZHEN DONSON CLOUD TECHNOLOGY CO LTD

Power distribution network tower defect automatic identification and classification method and system based on deep learning

The invention relates to a power distribution network tower defect automatic identification and classification method and system based on deep learning. According to the method, firstly, a tower main body area is positioned and extracted through a convolutional neural network, and background interference is eliminated; and then the visual saliency of the defect area is improved by adopting a self-adaptive contrast enhancement algorithm based on local statistical characteristics. In the feature extraction stage, a pyramid distraction attention module is introduced to fuse multi-scale space information and channel attention, and a two-dimensional selective state space module is used for modeling a long-range dependency relationship. And multi-resolution features are further aggregated through a layered feature fusion architecture and a self-adaptive anchor frame mechanism, and targets of different sizes are matched. And finally, a self-adaptive edge enhancement module is adopted to enhance the edge of the defect, and a multi-branch detection head is adopted to realize category judgment, position regression and confidence evaluation of the defect in parallel. The method effectively improves the detection precision and robustness of tower defects under a complex background, and is especially suitable for the automatic recognition of micro-scale defects.
Owner:SICHUAN YAAN ELECTRIC POWER (GRP) CO LTD +1

Video stream-based beef cattle feed intake intelligent monitoring and analysis method

The invention relates to the technical field of image analysis, in particular to a beef cattle feed intake intelligent monitoring and analysis method based on a video stream, which comprises the following steps of: converting a single-frame beef cattle house RGB image in the video stream into a logarithmic chroma space; according to the method, the chromaticity space conversion is performed on the single-frame cowshed RGB image and the illumination and reflection components are separated, so that the shadow interference can be suppressed, the feeding trough area is more stable in the image, the shadow suppression feeding trough image is divided into overlapped blocks, and local structure features and visual saliency features are calculated to generate a priority matrix; the local enhancement process can be adjusted according to actual feature distribution instead of a single global index, the flexibility and accuracy of the local area in contrast adjustment are improved, the self-adaptive contrast limit amplitude is set for different blocks in combination with the enhancement priority matrix, the distinction degree of the feed area and the background is improved, and the contrast adjustment accuracy is improved. And the texture and boundary features in the region are more prominent.
Owner:INST OF ANIMAL SCI & VETERINARY HUBEI ACADEMY OF AGRI SCI

Hyperspectral anomaly detection method and system for enhancing low rank and significance

The invention discloses a hyperspectral anomaly detection method and system for enhancing low rank and significance. The method comprises the following steps: acquiring hyperspectral image tensor data of a to-be-detected region; carrying out band-by-band normalization processing; decomposing into a background tensor and an abnormal tensor; designing an improved weighted tensor nuclear norm based on tensor singular value decomposition; constructing a visual saliency sparse weight tensor; establishing an anomaly detection model; obtaining an optimal abnormal tensor; and detecting the obtained optimal abnormal tensor to generate an abnormal detection result graph. According to the method, comprehensive mining of background low-rank features and full utilization of abnormal target saliency can be realized, and the detection rate of hyperspectral anomaly detection is effectively improved; the abnormal target can be clearly and accurately detected, the shape of the abnormal target can be completely reserved, good detection performance is still kept under the complex background and noise interference, and reliable technical support is provided for related applications such as mineral exploration and military reconnaissance.
Owner:SHAOXING UNIVERSITY

Display screen layout control method and system based on AI

The invention relates to the technical field of display screen layout control, and discloses an AI-based display screen layout control method and system, and the method comprises the steps: obtaining the operation record and focus distribution data of a user, and obtaining the time sequence data of a user behavior; extracting dynamic change rule characteristics of user behaviors according to the time sequence data; user interaction records are mined based on dynamic change rule features, classification is carried out according to time and scene dimensions, and long-term trend features and short-term fluctuation features of user behaviors are determined; obtaining the attention weight value of each layout element, and determining a visual saliency sequence; priority values of high-priority elements are calculated in combination with visual saliency sorting, and a preliminary layout scheme is obtained; performing amplitude adjustment according to the preliminary layout scheme to obtain an adjusted layout scheme, and adaptively changing the sizes and positions of layout elements to obtain an optimized display screen layout. According to the method, the display screen interface layout can be dynamically optimized, and the user interaction experience and the working efficiency are improved.
Owner:SHENZHEN FRIDA LCD CO LTD

Self-adaptive control method and system for cleaning unmanned aerial vehicle based on neural network

The invention provides a cleaning unmanned aerial vehicle self-adaptive control method and system based on a neural network, and relates to the technical field of cleaning. Multi-source sensing information of a cleaning unmanned aerial vehicle is acquired through a sensor assembly; activating a visual saliency perception center to process and perceive a to-be-cleaned target image in the multi-source sensing information to obtain a to-be-cleaned point location; generating a self-adaptive cleaning path in cooperation with unmanned aerial vehicle attitude data and cleaning environment data in the multi-source sensing information; and taking the to-be-cleaned coefficient of the to-be-cleaned point as a constraint, adjusting the initial cleaning strategy to obtain a target cleaning strategy, and performing cleaning adaptive control on the cleaning unmanned aerial vehicle in combination with the adaptive cleaning path. The technical problems of identification deviation, path rigidity and cleaning strategy imbalance of the cleaning unmanned aerial vehicle in a complex task environment due to lack of an effective information fusion and adaptive decision mechanism in the prior art are solved, and the technical effect of realizing precise and intelligent cleaning operation is achieved.
Owner:INNER MONGOLIA UNIV OF TECH

Video encoding method, video encoder, electronic device, and medium

PCT designated stageWO2026174439A1Pattern recognitionAcquisition apparatus
A video encoding method, a video encoder, an electronic device, and a medium, relating to the technical field of video processing. The video encoding method may comprise: acquiring a video stream to be encoded; for a current frame in the video stream, acquiring motion vectors and depth values respectively corresponding to a plurality of image blocks in the current frame; acquiring, on the basis of the depth values and the motion vectors, control parameters respectively corresponding to the plurality of image blocks, wherein each depth value represents a distance between content represented by the respective image block and a video acquisition apparatus, and each motion vector represents motion of the respective image block in the video stream; and encoding the current frame on the basis of the control parameters respectively corresponding to the plurality of image blocks, wherein the control parameters are used for representing visual saliency of the image blocks in the video stream, and are at least used for participating in discarding redundant information in the current frame during encoding.
Owner:BOE TECHNOLOGY GROUP CO LTD

Industrial image anomaly detection system based on double-branch prototype residual network

The invention discloses an industrial image anomaly detection system based on a double-branch prototype residual network. The system comprises an anomaly generator module, an improved ResNet-18 pre-training network module, a multi-scale fusion module, a multi-scale prototype sample library module and a multi-scale reversible self-attention module. Firstly, a new anomaly generator is put forward, diversity of anomaly samples is effectively expanded, the problems of data scarcity and imbalance are solved, secondly, a multi-scale prototype sample library is introduced in consideration that industrial defects have weak visual saliency and complex multi-modal features, anomaly detection precision and robustness are improved, and the anomaly detection accuracy is improved. And finally, based on an improved ResNet-18 network and a multi-scale reversible self-attention module, accurate positioning and information integrity of an abnormal region are ensured, and the detection rate of an anomaly detection task is improved. According to the method, a good effect can be achieved in the anomaly detection and positioning task of the industrial image, and the accuracy and recall rate of anomaly positioning can be remarkably improved.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Desktop type high-resolution multi-dimensional optical image intelligent analysis method and device

The invention relates to a desktop type high-resolution multi-dimensional optical image intelligent analysis method and device, and belongs to the technical field of optical image analysis. The method comprises the following steps: controlling an imaging device and an integrated desktop objective table through a host, performing synchronous image acquisition on a sample on the objective table, and generating an original multi-dimensional image set; generating a single-frame two-dimensional light and shadow mapping graph through a parallax fusion algorithm, and simulating the three-dimensional morphology characteristics of the sample through gray scale and color gradient change; outputting the dynamic light and shadow mapping graph to a desktop display screen in real time, and loading an interactive light source simulation engine; simulating an illumination effect on the dynamic light and shadow mapping graph in real time; the visual saliency of target features is improved by automatically configuring light source parameters; and according to the optimized rendering result, topological structure features and optical features of the sample are extracted, and a quantitative analysis report is generated. The full-process automation and intelligentization of optical images from multidirectional and multi-view collection to deep analysis and quantitative report generation are realized.
Owner:SHANGHAI HENGGUANG POLICE EQUIP

Pedestrian adversarial texture generation method and device based on double attention guidance

The present application belongs to the field of artificial intelligence and computer vision security technology, and provides a pedestrian adversarial texture generation method and device based on double attention guidance, which constructs scene prompt words through a large language model and combines a generative diffusion model to generate a base texture highly consistent with the current environment semantics before optimization, uses a three-dimensional visual prediction network to extract the geometric features of the grid and the generated initial texture feature map to construct a human eye visual saliency mask, generates a model sensitivity mask through a target detector, and based on a non-gradient color constraint mechanism of an HSV space, limits the texture within a physically achievable color space and suppresses high-frequency noise, solving the color distortion problem in physical implementation of traditional methods, and greatly improving visual concealment while maintaining strong attack robustness under multiple viewing angles and distances.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Screen image copyright protection method based on visual saliency multi-scale Hash

The invention discloses a screen image copyright protection method based on visual saliency multi-scale Hash, which is composed of a visual saliency enhancement module, a dual-path dynamic multi-scale feature extraction network and a CP tensor Hash layer. The method is characterized in that text / icon region features are enhanced through a significance model fusing edge density and color sparsity analysis; constructing a parallel network containing a detail path and a semantic path, and adaptively aggregating shallow detail and deep semantic features by using a cross-scale attention fusion mechanism; and mapping the high-dimensional feature into a compact binary hash code by adopting CP tensor decomposition. The method effectively solves the problems of insufficient capture of screen image structured elements and limited single-scale feature representation in the traditional technology, significantly improves the accuracy and robustness of copyright detection through the multi-granularity feature fusion and tensor compression technology, and is suitable for screen image copyright authentication scenes such as digital documents and UI interfaces.
Owner:GUANGXI NORMAL UNIV +1

Devices, methods, and graphical user interfaces for displaying movement of virtual objects in communication session

A computer system displays a representation of a user pose in a three-dimensional environment in response to movement of a user's current viewpoint. The computer system displays different representations of the movement of the virtual representation based on the type of the virtual representation of the user. A computer system reduces visual saliency of a virtual representation when changing a spatial arrangement of virtual objects shared in a communication session. The computer system displays different visual feedback when moving the virtual object depending on whether the virtual object is shared or not shared in the communication session. The computer system displays visual feedback indicative of audio provided by another user. The computer system displays feedback indicating that the participant will correspond to the location. The computer system displays a sequence of visual transitions while displaying a visual representation of a participant in the communication session.
Owner:APPLE INC

Remote Sensing Image Text Retrieval Method Based on Remote Sensing Multimodal Model

This invention relates to the fields of remote sensing image analysis and cross-modal retrieval technology. It discloses a remote sensing image text retrieval method based on a remote sensing multimodal fundamental model. By fine-tuning the CLIP large model, its large-scale pre-training capabilities are applied to the semantic alignment of remote sensing images and text. By introducing a visual saliency calculation module and a visual block fine-grained selection integration module, this invention effectively solves the problem of multi-scale targets and redundant information in remote sensing images, achieving fine-grained semantic alignment between images and text, and improving retrieval accuracy. Especially when images contain multiple salient targets and redundant regions, the proposed cross-modal semantic alignment fine-grained filtering method can accurately identify key information blocks in the image and perform fine matching with the text description.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Image encoding, decoding method and apparatus, codec

An image encoding and decoding method and apparatus, and an encoder and decoder, are disclosed. The method includes: acquiring a visual saliency heatmap of an image in the current frame, and filtering the current frame image using the visual saliency heatmap to obtain a target image; acquiring a motion estimation vector and a target prediction image of the next frame input image using the target image and the next frame input image; and encoding the difference image between the next frame input image and the target prediction image and the motion estimation vector. This improves the accuracy of acquiring the target prediction image and also improves the encoding accuracy.
Owner:BOE TECHNOLOGY GROUP CO LTD

Role animation key frame simplification method and system based on visual saliency

The invention discloses a character animation key frame simplification method and system based on visual saliency, and relates to the technical field of image processing, and the method comprises the steps: obtaining a character animation sequence, and extracting a position path and a rotation path of each skeleton node in the character animation sequence in a time dimension; calculating a first motion feature sequence based on the position path of each skeleton node; calculating a second motion feature sequence based on the rotation path of each skeleton node; combining the first motion feature sequence and the second motion feature sequence of each skeleton node into a third motion feature sequence; and selecting a target key frame sequence from the original time points based on the third motion feature sequences of all skeleton nodes by taking minimization of a motion reconstruction error of the role animation sequence as an optimization target, and outputting the target key frame sequence. Through feature extraction and fusion combination of the position path and the rotation path, accurate screening of animation key frames is realized, redundant frames are reduced, and data compression efficiency and animation coherence are improved.
Owner:GUANGZHOU LINKAGE NETWORK TECHNOLOGY CO LTD

Course video key frame intelligent identification method based on AI visual attention mechanism

The application relates to the technical field of video recognition, and discloses a course video key frame intelligent recognition method based on an AI visual attention mechanism, which comprises the following steps: acquiring a visual saliency feature map of a video frame and calculating a global attention gravity center coordinate, constructing a spatial second moment tensor by using the visual saliency feature map to determine an anisotropy coefficient, then performing nonlinear weighted processing on the trajectory distribution density in a space-time trajectory space, calculating a second acceleration residual based on the processed trajectory evolution process, and determining a key frame in combination with the trajectory distribution density and the second acceleration residual. The application uses an anisotropy regulation mechanism to suppress non-content dynamic interference, checks the integrity of teaching content generation through the second acceleration residual, solves the problem of lag in semantic turning point capture under a dynamic background, and enhances the semantic density of extracted sequences and the recognition stability.
Owner:HUNAN YUNPAN NETWORK TECH CO LTD

Graph coding and image decoding method of dot matrix anti-counterfeiting code

The invention relates to a graphic coding and image decoding method of a dot matrix anti-counterfeiting code, which optimizes the coordinate layout of code points by designing template coding units E-S and positioning structure code point combination rules in the aspect of graphic coding of the dot matrix anti-counterfeiting code. In combination with a corner sharing splicing extension method, a data region assignment strategy based on a dynamic mask and a boundary region assignment strategy based on randomization, the visual saliency of a positioning pattern is effectively reduced, texture traces during splicing printing are inhibited, and the visual quality of a coded image is remarkably improved; in the aspect of image decoding of dot matrix anti-counterfeiting codes, a UNet neural network image segmentation method based on deep learning and a spatial neighbor search algorithm based on a KDTree data structure are adopted, and accurate decoding of low-quality images and multi-unit complex dot matrix anti-counterfeiting codes is achieved. The method solves the problems of dominant exposure of a positioning structure, insufficient decoding robustness and the like, and has the characteristics of good concealment of the positioning structure, strong anti-copying capability, excellent visual effect and the like.
Owner:XIDIAN UNIV

A method for complex motion perception based on human visual cues

The application discloses a complex motion perception method based on human visual inspiration, which comprises the following steps: firstly, capturing the image of an athlete in a motion process, identifying a key image area by evaluating the semantics and visual importance of the image area, and dynamically processing the area features of the key image area; secondly, processing the dynamically processed area features, using a low-rank active learning technology, and obtaining an optimized image area through an optimized objective function; and finally, classifying the optimized image area by using a support vector machine (SVM) to obtain a motion perception result. The application can ensure better calculation efficiency and the visual saliency of the extracted key area, and significantly improves the performance in a complex motion environment.
Owner:HANGZHOU DIANZI UNIV

Image feature extraction-based topic token generation system and method

The application relates to the technical field of data processing, and discloses a subject Token generation system and method based on image feature extraction. The system integrates and divides semantic regions by hierarchically extracting and fusing multi-level visual features, obtaining a fused feature map, and providing basic feature data for subsequent semantic matching and feature quantization. According to domain subject knowledge, a subject semantic structure is built, and a quantization codebook is initialized. A mapping relationship between codebook entries and subject semantics is established, so that the codebook is self-adaptive to different image semantic scenes. In the attention interaction link, the mapping relationship is introduced to impose semantic constraints, and a sparse subject Token set is generated, thereby improving the authenticity and effectiveness of subject representation content. The saliency score of the subject Token is calculated by fusing the visual saliency and attention weight of the semantic region. After global deduplication and sorting, a subject Token sequence is formed, which comprehensively and objectively reflects the core subject information of the image.
Owner:SHANGHAI JIDOU TECH CO LTD

Adaptive dimming system for aviation obstruction light based on wireless networking

The present application relates to the field of intelligent control and photoelectric detection of aviation obstruction light, in particular to an aviation obstruction light adaptive dimming system based on wireless networking; the system comprises an environment sensing module, a feature coupling module, a control solving module and an adaptive driving module; the ambient light intensity value, the background environment color temperature value and the space scattering evaluation value are collected, the multi-dimensional light feature coupling compensation coefficient is solved, the background environment light comprehensive interference value is corrected and generated, and the target light source effective light intensity and its ratio are set as the dynamic ambient light signal-to-noise ratio; then the preset visual compensation model is input, the visual saliency maintenance index is solved and the dimming control instruction is generated, the driving current and the pulse duty cycle are adjusted, so that the visual contrast of the target light source is maintained within the preset safety interval, the present application can ensure the long-distance warning and recognition ability, suppress the near-distance glare and halo aggravation, and take into account the warning effect, thermal stability and device life.
Owner:CHENGDU JINHUA UTILITY ELECTRICAL APP RES INST