Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

106 results about "Transform coding" patented technology

Transform coding is a type of data compression for "natural" data like audio signals or photographic images. The transformation is typically lossless (perfectly reversible) on its own but is used to enable better (more targeted) quantization, which then results in a lower quality copy of the original input (lossy compression).

Traffic large model construction and decision-making method and device based on multi-modal two-way map reasoning

The invention discloses a traffic large model construction and decision-making method and device based on multi-modal two-way map reasoning, and the method comprises the steps: constructing a multi-modal data set of a text, an image and a track, generating fusion features through spatial-temporal clustering and cross-modal Transform coding, carrying out the two-way map reasoning in combination with a traffic knowledge map, and carrying out the decision-making of the traffic large model. The method comprises the following steps: generating an embedded representation through a forward graph neural network, reversely mapping a decision scheme generated by a language model to a graph to verify consistency, outputting knowledge to enhance embedding, fusing multi-modal features and knowledge embedding by adopting an LoRA multi-task joint fine tuning technology, adapting to traffic field tasks, deploying a real-time inference engine, and carrying out real-time inference on the traffic field. And processing the dynamic data flow through an aging perception attention mechanism, and outputting traffic event identification, path planning and scene question and answer results in parallel. Compared with the prior art, the method has the advantages that the problems of insufficient multi-source heterogeneous data fusion, low knowledge utilization efficiency and poor real-time decision consistency can be solved, and the semantic understanding and decision accuracy of the traffic large model is effectively improved.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

Mechanical arm path planning method based on Transform and diffusion model

The invention discloses a mechanical arm path planning method based on Transform and a diffusion model, and belongs to the technical field of robots, and the method comprises the following steps: collecting environment information to generate a reference path, and constructing a training data set; a conditional diffusion Transform prediction network is constructed, features are extracted, and multi-modal path prediction is realized; gaussian noise is applied to the reference path, and denoising training is carried out on the conditional diffusion Transform prediction network; in combination with a diffusion model and a cost guidance mechanism, optimizing a noise path sampled from Gaussian distribution until a smooth collision-free path is generated; the candidate paths are evaluated, and a mechanical arm joint control instruction is generated; and the mechanical arm executes the planning track and carries out real-time sensing and online re-planning. According to the method, environment perception, Transform coding and diffusion generation are organically combined, a plurality of feasible tracks which are smooth and capable of avoiding obstacles are rapidly generated in a complex obstacle scene, and the method has good generalization ability and can adapt to different scenes and dimension changes.
Owner:BEIJING UNIV OF TECH

Landslide risk assessment method based on extreme rainfall and geology coupling model

The invention discloses a landslide risk assessment method based on an extreme rainfall and geology coupling model, and relates to the technical field of geological disasters. Comprising the following steps: S1, constructing a three-dimensional probability density field of a fracture network and a non-Gaussian random field model of a permeability coefficient tensor; s2, setting a physical kernel layer according to the non-Gaussian random field model, setting a data driving layer through space-time Transform coding, and constructing a graph attention network model; s3, generating an adversarial network through physical information, constructing extreme rainfall coupling data, and updating the non-Gaussian permeability coefficient random field model according to the graph attention network model; and S4, acquiring an entropy generation rate according to the mechanical field data, the seepage field data and the temperature field data, and determining a risk level. Physical interpretability grading early warning of landslide risks is realized, and meanwhile, risk space distribution can be visually displayed through a sliding surface probability cloud picture, so that accurate decision support is provided for disaster prevention and control.
Owner:HUNAN INSTITUTE OF ENGINEERING

Adaptive Data Processing System with Real-Time Anomaly Detection and Self-Healing

A system and method for adaptive data processing combining compression and encryption. The system analyzes input data characteristics, compares probability distributions, and creates a transformation matrix to convert data into a dyadic distribution. It generates a main data stream of transformed data and a secondary stream of transformation information. The system dynamically selects and applies processing techniques, including transformation, encoding, compression, and encryption algorithms, based on analyzed characteristics and real-time performance metrics. It compresses the main data stream using Huffman coding and implements security measures to protect the output. A feedback loop monitors technique effectiveness, updates a knowledge base, and influences future selections. The system can operate in lossless, lossy, or modified lossless modes, adapting to different application requirements. This approach offers an efficient solution for scenarios where both data reduction and security are critical concerns.
Owner:ATOMBEAM TECH INC

Battery replacement robot target point cloud completion method based on dynamic graph convolution

The invention discloses a dynamic graph convolution-based target point cloud completion method for a battery replacement robot, and the method comprises the steps: 1, reconstructing a complete target fastener model from a multi-view image through SFM and MVS technologies, and obtaining complete point cloud data; 2, constructing an incomplete-complete point cloud pair as training data by using a geometric constraint-based adaptive cutting strategy; 3, multi-resolution point cloud processing is adopted, and feature extraction and fusion are carried out according to three-level resolution; 4, on the basis of DGCNN dynamic graph convolution, in combination with multi-stage Edge Conv dynamic edge convolution and a Transform coding module, local geometric feature capture and global relation perception are realized; 5, point cloud generation adopts a pyramid step-by-step refining method, geometric details are added to each layer based on a previous layer result, and the number of points is gradually expanded; and 6, the training process is guided through a multi-target joint loss function, and the point cloud quality is improved while the precision is ensured. The geometric structure integrity of the point cloud of the target fastener is improved, and the operation error of follow-up operation of the battery replacement robot is reduced.
Owner:SOUTHEAST UNIV

Hydropower station water level detection method and system

The invention discloses a hydropower station water level detection method and system. The method comprises the steps that an original water flow characteristic parameter set is obtained through multiple sets of Doppler effect radar flow measuring devices; performing segmentation extraction to obtain a water flow feature sub-data set; inputting a Transform network coding layer to generate a water flow feature coding vector; outputting a preliminary water level predicted value sequence through a decoding layer; a radar flow measurement model is called for correction to obtain a correction sequence; time sequence integration is carried out to obtain a water level change curve. The system comprises a multi-source Doppler radar signal acquisition unit, a spatial-temporal feature segmentation extraction unit, a Transform coding processing unit, a water level preliminary prediction unit, a radar model correction unit and a time sequence integration output unit. According to the method and system, multi-source parameters are deeply fused, the environmental adaptability is improved, the continuity and reliability of a detection result are enhanced, the water level of the hydropower station can be accurately monitored, and support is provided for scheduling decision and safety prevention and control.
Owner:GUODIAN DADUHE ZHENTOUBA HYDROPOWER CONSTR CO LTD

Visual defect detection method based on multi-level Transform

The invention discloses a visual defect detection method based on multilevel Transform. The method comprises the following steps: firstly, extracting multi-scale features of an input image through a pre-trained convolutional neural network, and performing hierarchical fusion; then, a multi-level Transform coding and decoding structure is used for carrying out deep reconstruction on the fusion features, and long-distance dependency relationships of different granularities are captured through hierarchical patch segmentation and a self-attention mechanism; meanwhile, introducing a standard deviation to estimate branch learning reconstruction uncertainty; and finally, a pixel-level anomaly score graph is generated by calculating a normalized residual error and adopting a multi-scale anomaly score aggregation strategy, so that accurate anomaly positioning and discrimination are realized. According to the method, multi-scale feature representation and global semantic modeling capability are fused, the accuracy, robustness and generalization capability of anomaly detection are effectively improved, only normal sample training is needed, and the method is suitable for complex scenes such as industrial visual detection and has a good application prospect.
Owner:SOUTHEAST UNIV

Transform coding based on matrix-based intra prediction

Devices, systems and methods for digital video coding, which includes matrix-based intra prediction methods for video coding, are described. In a representative aspect, a method for video processing includes performing a conversion between a current video block of a video and a bitstream representation of the current video block according to a rule, where the rule specifies a relationship between applicability of a matrix based intra prediction (MIP) mode or a transform mode during the conversion, where the MIP mode includes determining a prediction block of the current video block by performing, on previously coded samples of the video, a boundary downsampling operation, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation, and where the transform mode specifies use of a transform operation for the determining the prediction block for the current video block.
Owner:BYTEDANCE INC +1

Single-pixel multi-encryption imaging method and device and storage medium

The invention discloses a single-pixel multi-encryption imaging method and device and a storage medium. The method comprises the steps of generating a single-pixel orthogonal basis transform coding illumination pattern, and exchanging the sequence of a coding illumination pattern sequence by using a time sequence key and a remapping algorithm; the coded illumination pattern is processed by using a space key, and at the moment, an image sequence is encrypted in two aspects of a space domain and a frequency domain; projecting a preset coding illumination pattern to the surface of an object to be encrypted and imaged; measuring and recording the light intensity of the reflected or transmitted light by using a detector; decrypting the signal by using the combination of the space key and the time sequence key to obtain a plaintext signal; and reconstructing the image by using a corresponding single-pixel imaging method. According to the method, the single-pixel imaging technology is utilized, the image encryption signal can be obtained under the condition of low sampling rate, calculation only needs to be carried out when the encryption coding illumination pattern is generated, and the defect that a traditional image encryption method based on a camera is large in calculation amount is overcome.
Owner:HUZHOU COLLEGE

Flash furnace fault prediction method based on graph fusion and multi-stage learning

The invention relates to the technical field of fault prediction, and relates to a flash furnace fault prediction method based on graph fusion and multi-stage learning, which comprises the following steps: establishing a static graph and a dynamic graph, describing the influence intensity and dynamic change between sensor data according to the static graph and the dynamic graph, and obtaining spatial-temporal characteristics; inputting the spatial-temporal characteristics into a Transform coding model, and obtaining prediction data; the Transform coding model is trained by using a training set, and model parameters are adjusted by minimizing a prediction error to obtain the Transform coding model; the prediction data is coded, the coded data is stretched, translated and rotated, each data point is mapped to a high-dimensional space, the distribution position of abnormal points is obtained, the abnormal score of the data is calculated, the abnormal score is judged, and the fault result of the nickel flash furnace system is obtained. According to the method, the static relation and the dynamic relation can be mined and fused from various sensor data, so that fault prediction is realized.
Owner:LANZHOU UNIVERSITY OF TECHNOLOGY

Fashion preference prediction method and device

The invention relates to the technical field of multi-modal data prediction, and discloses a fashion preference prediction method and device, and the method comprises the steps: obtaining multi-modal data comprising an image, a text and a user behavior time sequence, carrying out the feature extraction, and obtaining an image feature, a text feature and a time sequence feature; constructing spatial features based on the city embedded table and the corresponding regional culture label codes; aligning the dimensions of the spatial features and the time sequence features, and fusing the aligned spatial features and time sequence features by using a gating attention mechanism to obtain space-time fusion features; splicing the space-time fusion feature with the image feature and the text feature to obtain a multi-modal fusion feature; decoupling the multi-modal fusion feature to obtain a material decoupling feature, a style decoupling feature and a scene decoupling feature; and after feature splicing is carried out on the multiple decoupling features, the fashion preference prediction probability is obtained through Transform coding and MLP classification.
Owner:SUZHOU UNIV

Monocular dressing human body reconstruction method for perspective distortion image

The invention discloses a perspective distortion image-oriented monocular dressing human body reconstruction method, and relates to the technical field of three-dimensional vision. The invention provides a three-dimensional human body reconstruction framework fusing distortion correction, three-dimensional geometric representation and pseudo multi-view constraint, and image distortion is relieved by uniformly partitioning an image and using virtual view transformation to decouple the relevance between an image position and perspective distortion; a multi-scale attention enhancement network is designed, and a hourglass model and a channel attention mechanism are utilized to extract block features, so that the model can process perspective distortion images; the human body parameter model is coded into a frequency domain coefficient through Fourier transform, and the frequency domain coefficient serves as the basis of geometric representation; in order to further optimize high-frequency details, the invention provides a pseudo multi-view module, and the robustness of the network to perspective distortion is enhanced by simulating multi-view data enhancement. Experimental results show that compared with the prior art, the method disclosed by the invention has the advantages that the quantification and the qualitative property on the Thuman2.0 data set and the CumHumans data set are optimal.
Owner:TIANJIN UNIV

Dynamic space-time CNN-Transform emotion brain-computer interface decoding method

The invention discloses a dynamic space-time CNN-Transform emotion brain-computer interface decoding method, and relates to the technical field of brain-computer interfaces, a dynamic time feature extraction module is designed according to the sensitivity of multi-scale convolution to an electroencephalogram sequence along a time dimension, and electroencephalogram sequence time feature information is mined by using convolution kernels of different sizes; constructing a local-global spatial feature extraction module by referring to close correlation between asymmetry of left and right brain regions of the brain and an emotional state, and sequentially extracting spatial feature information of the left brain, the right brain and the whole brain of the electroencephalogram sequence; a spatial-temporal feature fusion module is designed, and spatial-temporal feature relations among the left brain, the right brain and the whole brain are mined; the long-time dependency relationship in the electroencephalogram sequence is captured through an attention mechanism by referring to the advantage of Transform on long-time sequence processing, so that the emotion electroencephalogram decoding precision is effectively improved; and finally, performing emotion recognition on the feature sequence subjected to Transform coding by using a multi-layer perceptron to realize end-to-end emotion electroencephalogram decoding.
Owner:SHANGHAI UNIV

Gold mine target region optimization method and system based on spatial constraint modified embedded geochemical anomaly and memory

The invention discloses a gold mine target region optimization method and system based on spatial constraint modified embedded geochemical anomalies and a memory. The method comprises the following steps: acquiring geochemical data of a target area, and preprocessing to obtain a spatially serialized standardized data matrix and an adjacent matrix; according to the method, global Transform coding modeling is adopted to obtain a hidden state sequence of global element interaction characteristics of each sampling point of a target area, then unitary potential energy and binary potential energy are constructed in sequence, a geochemical abnormal area identification model of each sampling point of the target area is established in combination with CRF, and the geochemical abnormal area marked by the model is subjected to abnormal scoring and marking. And through comprehensive evaluation, gold mine target region optimization is completed. Based on the system containing the computer program memory provided by the method, the problems of large calculation amount and low accuracy in the prior art can be effectively solved, and through tests, the area under the ROC curve of the gold mine target region optimization process by adopting the system can reach 0.92, and the obvious advantages of the system are shown.
Owner:CENT SOUTH UNIV

High-precision remote sensing image cultivated land boundary extraction method

The invention discloses a high-precision remote sensing image cultivated land boundary extraction method, which comprises the following steps of: acquiring high-resolution remote sensing image data, preprocessing the data, and obtaining cultivated land parcel vector data through manual annotation; constructing a multi-task depth semantic segmentation model based on a coding and decoding structure; constructing a mixed loss function; acquiring a high-resolution remote sensing image of a cultivated land area to be detected and performing data preprocessing; inputting a high-resolution remote sensing image of a cultivated land area to be detected into the trained high-resolution remote sensing image cultivated land extraction model based on multi-task learning; and converting the cultivated land semantic segmentation result into cultivated land parcel vector data containing latitude and longitude coordinates. According to the method, through staged Transform coding based on overlapping patches and efficient sequence reduction, efficient multi-scale coding is realized, the problem of insufficient remote dependence capture is solved, and the computing power bottleneck of high-resolution self-attention is relieved.
Owner:HUZHOU CHUANGYI TECH CO LTD

Multi-modal data fusion and intelligent analysis method and system for advanced early warning of wind and light storage equipment

The invention discloses a multi-modal data fusion and intelligent analysis method and system for advanced early warning of wind and light storage equipment, and the method comprises the steps: deploying data collection equipment, synchronizing the data collection equipment to a virtual equipment model through a digital twin interface, and generating a physical enhancement feature; recombining the physical enhancement features into a three-dimensional feature tensor; orthogonal constraint tensor decomposition is carried out on the three-dimensional tensor, and cross-scale fault features are extracted; constructing a fault propagation graph network FPGN based on the core tensor, and outputting an early warning decision through graph convolution and Transform coding; according to the early warning deviation, a gradient descent method is adopted to adjust the physical constraint weight, and the model is retrained; the three-dimensional feature tensor recombination realizes unified characterization of multi-modal data, orthogonal constraint tensor decomposition forces different dimension features to be independent, a cross-scale coupling mode of a fault is extracted, and the effects of dimension reduction and efficiency improvement are achieved; according to the FPGA network, graph convolution and Transform time sequence coding are fused, and the propagation intensity of modeling faults in an equipment topology network is realized, so that early warning from local anomalies to system-level risks is realized.
Owner:SHAANXI HYDROPOWER DEVELOPMENT GROUP CO LTD +2

Verification of perception systems

ActiveUS12547879B2Neural learning methodsKnowledge based modelsAlgebraic transformationsAlgorithm
There is provided a computer-implemented method for verifying the robustness of a neural network classifier with respect to one or more parameterised transformations applied to an input, the classifier comprising one or more convolutional layers, the method comprising: encoding each layer of the classifier as one or more algebraic classifier constraints; encoding each transformation as one or more algebraic transformation constraints; encoding a change in an output classifier label from the classifier as an algebraic output constraint; determining whether a solution exists which satisfies the classifier constraints, transformation constraints and output constraints, and determining the classifier as robust to the local transformations if no such solution exists. A perception system and a computer readable medium are also provided.
Owner:IMPERIAL COLLEGE INNVOATIONS LTD

Knowledge graph multi-mode document analysis and image table semantization knowledge recall method

The invention discloses a knowledge graph multi-modal document analysis and image table semantization knowledge recall method, and belongs to the technical field of knowledge engineering and information retrieval. The invention provides an innovative scheme for fusing a visual language model, semantic abstract generation and knowledge graph modeling. The method comprises the following steps: constructing a vertical domain knowledge graph by adopting a BERT-BiLSTM-CRF model; according to the method, multi-modal document analysis is realized through models such as DocLayout-YOLO, TableMaster, UniMERNet and the like; the method comprises the following steps of: segmenting an image into 16 * 16 block sequences by adopting a vit-gpt2-image-adaptation model, and realizing image semantization through 768-dimensional vector space mapping and Transform coding; constructing a document summary tree based on DBSCAN clustering and LLM recursive summary; and designing a hybrid retrieval space fusing semantic vectors and structured vectors, and reordering by adopting a double-attention mechanism. According to the method, the knowledge base document retrieval recall rate is increased to 99%, the question and answer accuracy rate reaches 90% or above, the index construction time is shortened by 60%, and the problem that semantic understanding and recall of non-text elements in complex documents are difficult is effectively solved.
Owner:云鼎科技股份有限公司

Text classification method and device and processor

The invention discloses a text classification method and device and a processor. A normalized text vector is obtained by preprocessing and coding an input text. Then utilizing a multi-layer structure in a pre-trained bad text recognition model, firstly carrying out first-level category classification on the input vector through a first Transform coding layer and a classification layer, and meanwhile, carrying out related semantic clustering through a first clustering layer to generate a first clustering text vector; then, a second Transform coding layer and a classification layer complete finer-grained secondary category classification based on the first clustering text vector, and continuous aggregation is performed through a second clustering layer to generate a second clustering text vector; and finally, the third Transform coding layer and the classification layer are combined with the second clustering vector to output a global classification result. And finally, all classification results are screened by setting a confidence threshold, high-confidence categories are reserved, and accurate classification results with clear levels are obtained.
Owner:NEUSOFT CORP

A method, an apparatus and a computer program product for video encoding and decoding

A method comprising: processing a video frame; determining a reference block for a current block of the video frame; predicting the current block with an intra block copy method; deriving intra prediction information for the current block based on the reference block; selecting a transform coding for the current block based on the intra prediction information; and applying the selected transform coding to the current block.
Owner:NOKIA TECHNOLOGIES OY

Intelligent popular science content retrieval interaction system and method combined with knowledge graph

The invention discloses a science popularization content intelligent retrieval interaction system and method combined with a knowledge graph. The system comprises a knowledge graph construction unit, a multi-layer multi-head attention mechanism collaborative Transform coding unit, a knowledge graph embedding and Transform feature fusion unit, a retrieval result sorting unit based on fusion features, an interactive intention capture unit and a result optimization output unit. The method comprises the following steps: constructing a knowledge graph for multi-source science popularization data, extracting retrieval keyword features by applying Transform coding, fusing the knowledge graph and the retrieval keyword features, and calculating the similarity with science popularization content to finish initial sorting; and meanwhile, capturing user interaction behavior analysis intention change in real time, and readjusting and outputting an initial result in combination with knowledge graph semantic association. The method corresponds to operation steps of all units of the system, and intelligent retrieval interaction of popular science content is achieved. According to the system and the method, the problems of insufficient semantic understanding, poor interaction experience and the like of traditional retrieval are effectively solved, and the accuracy and the interaction intelligence of popular science content retrieval are improved.
Owner:SHENGDI XINGTU INFORMATION TECH CO LTD OF LHASA ECONOMIC & TECH DEV ZONE

360-degree image compression method and device using adaptive latitude-aware transform coding

The application discloses a 360-degree image compression method and device based on adaptive latitude-aware transform coding, has obvious advantages in code rate saving, and can effectively solve the distortion redundancy problem of ERP images.The method comprises the following steps: (1) designing an adaptive latitude-aware module; (2) constructing a multi-scale gated convolutional neural network; (3) guiding spatial features by transform modulation importance feature activation maps; and (4) constructing a learned overall framework of 360-degree image compression.
Owner:BEIJING UNIV OF TECH

A data asset management method, device, storage medium and system

The present application relates to the technical field of data processing, and especially relates to a data asset management method, device, storage medium and system, the method comprising: acquiring video data to be processed; determining key texture features and local chroma features of each video frame; cutting the video data to be processed into several video segments; determining a predicted compression interference characteristic value for each video segment based on the similarity of the key texture features of adjacent video frames in the video segment and the dispersion of the local chroma features of each video frame; configuring a label for each video segment according to the predicted compression interference characteristic value corresponding to the video segment; determining a division area for each video block in each region based on the spatial position relationship between the key texture features and each region in the video frame of the video data to be processed; performing transform coding on each residual value, or adjusting the number of predicted frames in each video segment based on the predicted compression interference characteristic value, and performing inter-frame prediction compression on each video frame, thereby improving the video compression efficiency and spatial utilization.
Owner:SEVEN (BEIJING) EDUCATION TECH CO LTD

Frequency-time fusion human motion prediction method

The invention discloses a frequency-time fusion human motion prediction method, and relates to the technical field of 3D human motion prediction. According to the method, the time domain representation and the frequency domain representation of the 3D human body motion are utilized at the same time for the first time to make up for the limitation brought by discrete cosine transform coding, on this basis, a novel model HIFT is provided, and the model can harmoniously integrate the frequency characteristics and the time domain characteristics, so that the prediction accuracy and reliability are improved; two modules, namely a frequency feature extraction module (FFM) and a time domain feature extraction module (TFM), are provided in an HIFT model, the frequency feature learning module FFM is formed by stacking multiple layers of perceptron, the time feature learning module TFM is constructed on the basis of a graph convolutional network, and in order to adaptively capture complex spatial dependence, the frequency feature extraction module FFM and the time domain feature extraction module TFM are used for adaptively capturing complex spatial dependence. The TFM provided by the invention is a double-graph dynamic convolution module, that is, the design of a double-graph convolution structure is adopted to better capture spatial features from time domain information, so that the defect of frequency domain information after DCT coding is made up.
Owner:CHONGQING UNIV OF TECH

Real-time decoding positioning method for coded targets of a gap measurement system

The present application belongs to the technical field of target real-time decoding positioning, and provides a kind of encoding target real-time decoding positioning method of slit measurement system, its steps include: S1 the layout of encoding target of slit detection point;S2 for the binary coding of encoding target image shot at large inclination angle, and the center position of encoding area and positioning area is obtained as the initial point of compression transformation;S3 adopts single direction wide edge detection to combine edge projection method, and is compressed to project transformation to encoding target point;S4 encoding target decoding;S5 subpixel positioning of encoding target edge pixel optimization, the present application provides edge projection benchmark using single direction wide edge detection, reduces the influence of compression projection transformation by large inclination angle shooting, realizes the compression target point projection transformation using edge projection, and the edge of positioning area is combined with the edge of encoding target, reduces the edge mis-extraction rate, improves the subpixel positioning precision, while having the characteristics of fast and real-time, suitable for long slit multi-point detection.
Owner:AVIC INTELLIGENT MEASUREMENT

Image coding method and device, equipment and storage medium

PendingCN121985144AImprove processing throughputReasonable workload distributionDigital video signal modificationComputer hardwareAlgorithm
The invention provides an image coding method and device, equipment and a storage medium, and relates to the technical field of computers. The method comprises the following steps that: a CPU (Central Processing Unit) analyzes an image to be coded in a wavelet transform coding process to obtain original pixel data, and transmits the original pixel data to a GPU (Graphics Processing Unit); the GPU performs preprocessing operation on the original pixel data in parallel by using a plurality of threads to obtain standard pixel data meeting the wavelet transform requirement; wherein one thread processes one pixel column in the original pixel data; the GPU performs forward wavelet transform on the standard pixel data to obtain a sub-band coefficient; and the GPU performs entropy coding processing on the sub-band coefficients to obtain coded data. According to the scheme, image analysis is completed in the CPU, and preprocessing, wavelet transform and entropy coding processing are executed in parallel in the GPU, so that the coding throughput rate and the resource utilization rate in the image coding process can be improved.
Owner:MOORE THREADS TECH CO LTD

Cerebrovascular segmentation method and device based on physical guidance and pyramid vision Transform

The invention discloses a cerebrovascular segmentation method and device based on physical guidance and pyramid vision Transform, and relates to the field of medical image data, and the method comprises the steps: constructing a cerebrovascular segmentation model, and enabling loss functions used during training to comprise boundary intersection-to-union ratio loss, focus Tversky loss and Dice loss; the method comprises the following steps: acquiring an optical coherence tomography image of a brain to be processed, inputting the optical coherence tomography image into a trained cerebrovascular segmentation model, enabling the optical coherence tomography image to pass through an encoder module of pyramid vision Transform, and inputting an output feature of a first Transform encoding layer into a radial strength module to obtain a radial enhancement feature; wherein the output features of the second Transform coding layer, the third Transform coding layer and the fourth Transform coding layer are input into a deformable cross-scale fusion module to obtain enhanced fusion features, and the radial enhanced features and the enhanced fusion features are input into a boundary perception attention module to obtain a corresponding cerebrovascular prediction segmentation mask and a cerebrovascular prediction segmentation image. The problems of low segmentation accuracy and boundary precision in the prior art are solved.
Owner:THE FIRST AFFILIATED HOSPITAL OF XIAMEN UNIV +1

Xinjiang cold and arid slope risk prevention and control method based on knowledge graph

PendingCN122656360AAridKnowledge graph
The application discloses a Xinjiang cold and arid slope risk prevention and control method based on a knowledge graph, belongs to the technical field of cold and arid region slope disaster risk prevention and control, and aims at solving the problems of non-uniform data format, weak correlation and insufficient prevention and control measures matching of cold and arid slope disasters. The application realizes the technical effects of improving the risk chain identification accuracy and the measure matching by constructing a heterogeneous space-time knowledge graph, generating multiple types of vectors, adopting a space-time graph transformation coding model to identify the instability chain and search similar cases, outputting a risk level, a treatment priority and recommended measures.
Owner:HOHAI UNIV

Video text retrieval method and system based on text condition semantics

PendingCN122654360ATime domainFrame sequence
The application provides a video text retrieval method and system based on text condition semantics, and relates to the technical field of natural language processing.The method comprises the following steps: according to a video frame sequence to be retrieved and a query text, performing feature extraction through a pre-trained visual-linguistic double-branch feature extraction network to obtain a frame-level visual feature sequence and a global text feature; performing scale sampling according to a plurality of preset time steps based on the frame-level visual feature sequence, and performing time domain transformation coding respectively to obtain a plurality of groups of time sequence features corresponding to different time granularities; after the plurality of groups of time sequence features corresponding to different time granularities are up-sampled and aligned to a unified time resolution, the integral of each time granularity of the time sequence features under a preset fractional order is calculated respectively to obtain multi-scale historical cumulative features.The application improves the accuracy of video text retrieval and the rationality of related result sorting.
Owner:HANGZHOU DIANZI UNIV