Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1129 results about "Coding decoding" patented technology

Edge-deployed semi-supervised anomaly detection method and system for railway track foreign object

Disclosed in the present invention are an edge-deployed semi-supervised anomaly detection method and system for a railway track foreign object. The method comprises the following steps: an edge device encoding and decoding a video stream captured by a camera to obtain an image frame sequence, and performing frame extraction; and using a semantic segmentation model to perform image segmentation on a certain image frame obtained by means of frame extraction, to obtain a railway track region segmentation image. The use of a single image as input may generate an expert model result having a high weight value; however, the determination based on a single image is not stable, multiple consecutive images of the task scene need to be inputted, the frequency of each expert model obtaining the highest weight is computed, and the expert model corresponding to the highest frequency is the final solution. The present invention supports scene-adaptive foreign object detection algorithm automatic selection, and a user can perform selection on the basis of prior knowledge, or selection may be performed by a scene-adaptive automatic algorithm selection method; the user only needs to provide a batch of image data of the current scene, and the optimal algorithm selection can be evaluated.
Owner:GUANGZHOU EMBEDDED MACHINE TECH CO LTD

Method, device, and recording medium for image encoding / decoding

Disclosed herein are a method, an apparatus and a storage medium for image encoding / decoding. In typical image encoding / decoding methods, a decoder-side motion information derivation method may be limitedly used. Therefore, the improvement of encoding efficiency attributable to the decoder-side motion information derivation method may also be limited. In embodiments, a motion information search method used in an inter-prediction mode, an intra block copy mode and an intra template matching prediction mode is disclosed. With the use of various motion search methods, encoding efficiency may be improved.
Owner:ELECTRONICS & TELECOMM RES INST

Method, device, and recording medium for image encoding / decoding

Disclosed herein are a method, an apparatus and a storage medium for image encoding / decoding. Methods for performing more accurate prediction are disclosed when performing prediction using a block vector. A block vector for a target block is derived by template matching, and a prediction block for the target block is determined based on the derived block vector. Various and detailed embodiments for a plurality of pieces of information used to derive a block vector, such as a search area, a block vector candidate list, a reference region, a buffer, and a tree type, are provided.
Owner:ELECTRONICS & TELECOMM RES INST

System and method for robot planning using large language models

A robotic controller for controlling a robot according to a sequence of robotic actions. comprises an input interface configured to receive a plurality of multimodal inputs each specifying instructions for performing a task in a different modality including audio, video, and a text modality. The controller also comprises a multimodal large language model, an action sequence decoder, and a controller. The multimodal LLM includes a multimodal LLM encoder and an LLM decoder. The multimodal LLM encoder is trained with machine learning to transform the multimodal instructions into encodings and the LLM decoder is configured to decode the encodings into a sequence of robotic instructions. The action sequence decoder is trained with machine learning to transform the sequence of robotic instructions into a sequence of actions using a library of robotic skills. The controller is configured to control a robot according to the sequence of actions.
Owner:MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC

Lightning approaching prediction method and device based on multi-source meteorological data

The invention discloses a thunder approaching prediction method and device based on multi-source meteorological data, and particularly relates to the technical field of thunder disaster prediction.The method comprises the steps that standardization processing of temporal-spatial resolution unification is conducted on meteorological satellite data, radar data and lightning positioning data, and standardized temporal-spatial input data is generated; then spatial-temporal features are extracted through a depth separable 3D convolution module, and a compressed spatial-temporal feature graph is generated; then, a channel-space double attention mechanism (CPCA) is applied to the compressed feature map for feature optimization; and finally, processing the optimized feature map through a coding-decoding structure, and outputting a thunder and lightning probability distribution map. According to the method, the problem of spatial-temporal resolution difference and physical feature mismatching in multi-source data fusion is innovatively solved, the calculation efficiency is remarkably improved through a lightweight network architecture, the feature expression ability is enhanced by using an attention mechanism, and high-precision prediction of sudden thunder and lightning events is realized.
Owner:CHENGDU UNIV OF INFORMATION TECH

Image defogging method based on dynamic wavelet prior and double-domain learning

The invention belongs to the technical field of image processing and deep learning, and particularly relates to an image defogging method based on dynamic wavelet prior and double-domain learning. Aiming at the requirements of all-weather clear imaging in the fields of intelligent traffic systems, safety monitoring and the like, and in order to overcome the defect that a static convolution kernel adopted by a traditional defogging method is difficult to adapt to different haze degradation, the invention provides a method for dynamically generating a convolution kernel by using haze priori contained in a multi-scale wavelet LL sub-band; and an efficient, robust and accurate image defogging model is constructed. According to the invention, based on a multi-scale U-shaped coding-decoding architecture, a dynamic wavelet depth separable convolution module DyWConv is embedded in front of each level of a coder to realize content adaptive feature extraction, and a double-domain feature learning module SPAFormer Block cooperatively utilizing Fourier domain global modulation and wavelet domain multi-scale decomposition is designed. And double-domain features are fully fused through an adaptive gating fusion mechanism, and finally a clear image is reconstructed and output step by step. According to the method, a method for explicitly encoding frequency domain degradation prior into dynamic convolution kernel parameters is innovatively provided, the complementary advantages of Fourier transform and wavelet transform are cooperatively utilized, spatial non-uniform haze can be effectively removed, image details can be recovered, leading performance is achieved in a synthetic data set and a real scene, and the method has a wide application prospect.
Owner:NANKAI UNIV

Multiplying active code decoding method based on SO-ORBGRAND decoder

The invention belongs to the technical field of communication, and particularly relates to a multiplication active code decoding method based on an SO-ORBGRAND decoder. The SO-ORBGRAND decoder outputs soft information by calculating a correct posterior probability of a conjecture code word sequence, and high-performance decoding is realized through soft information exchange and optimization processing between the row decoder and the column decoder. In the iteration process, a method for judging the validity of the code word through the check matrix is designed by utilizing the polarization characteristic of the multiplication active code, so that the iteration is supported to jump out in advance to reduce the complexity. The invention provides a multiplication active code decoding algorithm based on an SO-ORBGRAND decoder, and compared with various non-system multiplication active code decoding algorithms and the performance of traditional long-code polarization codes, the method has remarkable advantages in error correction performance; a method for judging the validity of the code word through the check matrix by utilizing the polarization characteristic design of the multiplication positive code is provided, the complexity is reduced while the judgment success rate is improved, and the implementation overhead is reduced.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Method and apparatus for encoding / decoding image

To provide a video encoding / decoding method for performing intra-prediction on the basis of availability of a reference pixel so as to improve encoding performance.SOLUTION: An intra-prediction method for a color copy mode checks a reference pixel region designated to obtain correlation information, determines a reference pixel processing setting based on a determination of availability of the reference pixel region, and performs intra-prediction according to the determined reference pixel processing. In addition, a candidate list for motion information prediction of the current block is generated, a control point vector of the current block is derived based on the candidate list and a candidate index, a motion vector of the current block is derived based on the control point vector of the current block, and inter prediction is performed on the current block.SELECTED DRAWING: Figure 14
Owner:INST OF IMAGE TECH INC

Image text description generation method, electronic equipment and readable storage medium

The invention provides an image text description generation method, electronic equipment and a readable storage medium. According to the method, the object perception prototype learning module and the global context feature extraction module are introduced, so that fine-grained information and global semantic understanding in the image are effectively balanced. The visual backbone network module can extract multi-scale and multi-level image features and perform fusion, thereby enhancing the expression ability of the image features. The object perception prototype learning module further extracts an object prototype from the fusion features to ensure that the model can accurately capture key objects and attributes thereof in the image, and the global context feature extraction module ensures that the overall context of the image is fully understood. On the basis, the encoding and decoding module combines the global context and the object prototype to generate the text description, so that the semantic splitting phenomenon in the traditional method is avoided, and the detail information in the image is effectively reserved, thereby improving the accuracy and integrity of the image description.
Owner:WUHAN UNIV

Coagulant addition prediction method based on fusion of sparse coding and graph space-time attention

The invention discloses a sparse coding and graph space-time attention fused coagulant addition prediction method, and relates to the technical field of time sequence data modeling prediction. According to the method, through the combination of a Transform encoding-decoding architecture and a sparse self-attention and top-k sparse strategy, the missing data complementation precision is effectively improved, and high-quality and missing-free input data is provided for subsequent prediction modeling; in the aspect of spatial feature expression, node association of multiple subsystems of a water production system is captured by means of a graph attention network, and by dynamically calculating association weights among nodes and aggregating neighbor features, the information entropy of spatial features is improved, and the expression ability of the spatial features to a multi-node coupling relationship is greatly enhanced; in the aspect of time sequence prediction performance, a time attention mechanism can dynamically focus on key time sequence nodes, the utilization rate of key time sequence features is further improved, and the accuracy of coagulant adding prediction is improved.
Owner:CHENGDU QIANJIA TECH CO LTD

Finite geometric code decoding method and system based on generalized check matrix

PendingCN120880462AAlgebraic geometric codesPhase-modulated carrier systemsTheoretical computer scienceBpsk modulation
The invention discloses a finite geometric code decoding method and system based on a generalized check matrix, and relates to the communication technology, and the method comprises the following steps: constructing a corresponding generalized check matrix according to an algebraic structure of a multi-step large number logic decodable finite geometric code; sending a finite geometric code sequence, transmitting the sequence after coding and BPSK modulation, and receiving a posterior probability log-likelihood ratio sequence by a receiving end; and carrying out iterative decoding on the finite geometric code received by the receiving end based on the constructed generalized check matrix. According to the invention, a new generalized check matrix is constructed according to the algebraic structure of the finite geometric code, and multi-step large-number logic iterative decoding is carried out on the finite geometric code based on the matrix, so that the iterative decoding performance of the multi-step large-number logic decodable finite geometric code is improved.
Owner:THE 20TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORP

Image encoding / decoding method and apparatus, and recording medium storing bitstream

This disclosure provides an image encoding / decoding method and device and recording medium storing a bitstream. The image decoding method may comprise: constructing a candidate list related to prediction of a current block; generating a prediction block of the current block based on a candidate selected from the candidate list; and generating a reconstructed block of the current block based on the prediction block, wherein the constructing of the candidate list includes: generating a candidate list for each of two or more types classified based on at least one coding parameter; and generating an integrated candidate list by adding at least one candidate selected in the candidate list for each type in a predetermined order, and wherein the at least one coding parameter is a motion information derivation method, and the two or more types include a first type of a spatially neighboring block, a second type of a temporal neighboring block and a third type of a non-neighboring block.
Owner:ELECTRONICS & TELECOMM RES INST

Learning-based point cloud geometry compression framework

In one implementation, geometry of a point cloud is encoded / decoded. On the encoder side, the encoder determines a first feature representing a voxel occupancy status of a current level and / or one or more finer levels of the point cloud, based on the voxel occupancy status of the current level and / or the finer levels; determines a second feature representing prediction of the voxel occupancy status of the current level and / or the finer levels of the point cloud; determines a third feature associated with the voxel occupancy status of the current level and / or the finer levels, based on the first feature and the second feature; and encodes the third feature. On the decoder side, the first feature is decoded from a bitstream, the second feature is determined similarly as the encoder side, and the third feature is determined based on the first feature and the second feature.
Owner:INTERDIGITAL VC HOLDINGS INC

Time sequence self-supervised learning method and system for rail transit engineering video images

The invention discloses a time sequence self-supervised learning method and system for a rail transit engineering video image, and the method comprises the steps: collecting an urban rail transit engineering construction video, and carrying out the segmentation of the video, and generating segmentation masks, so as to construct a self-supervised pre-training data set; a pre-trained model is selected and initialized; carrying out dynamic frame rate sampling processing on the input video; carrying out random shielding processing on each video frame to generate a corresponding binary mask image; carrying out multi-scale down-sampling on the processed video, and inputting the processed video into a coding-decoding network to complete spatial-temporal feature coding and generation of a reconstructed video and a prediction mask; a loss function is formed by calculating reconstruction loss of a reconstructed video and a real video, dynamic frequency loss based on Fourier domain low-frequency and high-frequency feature consistency and segmentation loss of a prediction mask and a real mask. The problems of insufficient time sequence modeling adaptability and insufficient multi-scale target adaptability are solved.
Owner:BEIJING URBAN CONSTRUCTION DESIGN & DEVELOPMENT GROUP CO LIMITED +1

Lightweight short-time rainfall prediction method based on context attention fusion

The invention discloses a lightweight short-time rainfall prediction method based on context attention fusion, and the method introduces multiple structural improvements on the basis of UNet, and comprises the steps: enabling a deformable convolution module to enhance the extraction capability of local complex features; the context expansion perception module improves the capability of capturing multi-scale features; the adjacent layer feature fusion module realizes information flow and detail retention capability among different levels of features; the SMamba module improves the long-range dependence modeling capability of the model, enhances the time sequence understanding of the model, introduces a depth separable convolution and jump connection mechanism in the encoding-decoding process, reduces the parameter quantity of the model, and improves the training stability of the model. According to the meteorological radar image sequence prediction method, the model parameter quantity and the calculation overhead are remarkably reduced on the premise of keeping the prediction precision, so that the model is more suitable for an application environment with higher requirements on the real-time performance and the resource consumption, and meanwhile, the model has good performance in a meteorological radar image sequence prediction task and has better application potential.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Low-illumination image enhancement method based on multi-mode classification and brightness feedback

The invention relates to a low-illumination image enhancement method based on multi-mode classification and brightness feedback, and belongs to the field of image processing. The method comprises the following steps of: firstly, training a network, constructing a multi-modal illumination prior feature of an image, inputting the multi-modal feature and a brightness score of the image into an illumination perception classification network, obtaining a local probability value of image brightness, adaptively selecting a local or global enhancement processing method, and generating a fusion weight; in the local enhancement processing, dark area details are enhanced through a multi-scale parallel and double-attention mechanism, in the global enhancement, the brightness is improved by using a symmetric coding-decoding structure and a residual attention block, and linear fusion is performed after the brightness is enhanced; then, the brightness evaluation value of the image is fed back through the lightweight brightness estimation network, and the brightness of the image is enhanced. And optimizing the training network according to the joint loss function of image processing. The method provided by the invention can enhance image texture details and improve image definition under a non-uniform illumination condition and an extremely low illumination condition.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Active learning and augmentation method for complex skill in heterogeneous scenario and sim-to-real transfer method

Provided in the present application are an active learning and augmentation method for a complex skill in a heterogeneous scenario and a sim-to-real transfer method. The active learning and augmentation method comprises: first, collecting state data by means of the interaction of multiple executors and an environment; storing same in a shared experience replay buffer; after sampling, alternately training the executors and discriminators; and finally selecting the optimal executor to be deployed, so as to implement active learning and augmentation of a complex skill. The sim-to-real transfer method comprises: in a 3C assembly digital twin environment, collecting multi-modal data to construct a skill knowledge base; generating a policy sequence; using technologies such as residual reinforcement learning to perform optimization; and, by means of a coding and decoding model, transferring a skill policy from simulation to reality. Without a teacher model or teaching data, the solution provided in the present application improves skill learning efficiency by means of combining reinforcement learning and knowledge distillation techniques, and implements the sim-to-real transfer of operating skills from a simulation environment to a real assembly environment.
Owner:TSINGHUA UNIVERSITY

Image codec method, image decode method, and decoder

To provide an image codec method, an image decode method and a decoder for effectively improving the efficiency of a codec.SOLUTION: Before performing the coding processing according to the MIP mode, the encoder performs a correction to unify initial right-shift parameters corresponding to different sizes and different MIP mode numbers according to an offset parameter for indicating the number of right-shift bits of the predicted value, and when performing the coding processing according to the MIP mode, the encoder performs the coding processing according to the offset parameter. The decoder may perform a correction to unify the initial right-shift parameters corresponding to different sizes and different MIP mode numbers according to the offset parameter before performing the decoding process according to the MIP mode, and perform the decoding process according to the offset parameter when performing the decoding process according to the MIP mode.SELECTED DRAWING: Figure 7
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Super-resolution structured light reconstruction method based on wavelet enhanced CNN-Transformer structure

The invention discloses a super-resolution structured light reconstruction method based on a wavelet enhanced CNN-Transformer structure. The super-resolution structured light reconstruction method comprises the following steps of: firstly, carrying out wavelet enhancement on a CNN-Transformer structure; according to the method, image reconstruction is realized based on a WECT network adopting a coding-decoding architecture; the method comprises the following steps: firstly, extracting initial features through shallow convolution, performing two-stage down-sampling by using discrete wavelet transform, and respectively inputting results into a convolutional neural network and a Swin Transform module to perform parallel feature extraction and fusion; secondly, performing two-stage up-sampling through inverse discrete wavelet transform, and performing jump connection with shallow layer features; finally, pixel shuffling up-sampling operation is introduced, and residual jump connection between the input image and the output image is established; according to the method, high-resolution details can be efficiently recovered while the structure reduction precision is ensured, the image definition and contrast are enhanced, and the application value and practicability of the SIM image in scenes such as biological imaging and tissue analysis are improved.
Owner:FUDAN UNIVERSITY

Multimodal ultrasonic microscopic image contrast enhancement method fusing expert priori knowledge

The invention discloses a multi-mode ultrasonic microscopic image contrast enhancement method fused with expert priori knowledge, and the contrast, definition and detection reliability of an ultrasonic microscopic image are improved. The method comprises the following steps: a multi-modal feature coding module based on a CLIP framework extracts feature representations of an ultrasonic microscopic image and an expert cue word, and maps the feature representations to a unified semantic space through a cross-modal alignment mechanism; designing a semantic guidance prompt module, constructing semantic elements which highlight detection demand guidance, and combining an image-quality description sample to carry out few-sample fine tuning so as to enhance the response capability of the model to specific semantics; an image enhancement and reconstruction module is designed, a coding-decoding structure and a cross-layer feature connection mechanism are adopted, contrast enhancement, detail reconstruction and noise suppression are realized under the guidance of a CLIP semantic vector, self-learning correction is carried out based on a standard grooving plate to improve contrast performance, and optimization training is carried out in combination with structural similarity loss and semantic consistency loss.
Owner:BEIJING UNIV OF CHEM TECH

Frequency modulation and wavelet sub-band guided double-domain cooperative Transform X-ray image denoising method

The invention discloses a frequency modulation and wavelet sub-band guided double-domain collaborative Transformer X-ray image denoising method, which comprises the following steps of: acquiring a noise-containing digital ray original image and a corresponding clear reference image, and constructing a data set after preprocessing the noise-containing digital ray original image and the corresponding clear reference image; constructing a network model of a double-domain collaborative coding-decoding architecture; performing 3 * 3 deep convolution on an input image to extract shallow layer features; in the encoding stage, ETB and AFMB are alternately stacked to represent local and global information, a WB-LKED module is embedded to strengthen fine-grained features, and WDB executes down-sampling and transmits high-frequency features to a decoding end; in the decoding stage, the WUB recovers the resolution through double-path up-sampling, integrates the same-scale features of an encoder, enhances details by using high-frequency features, splices the features, then carries out ETB and AFMB refining, obtains output features through 3 * 3 deep convolution, and combines a global residual error connection optimization result; and training the model by using the data set, inputting a to-be-denoised image, and outputting a final result. According to the method, the problems of weak complex noise interference resistance, poor detail retention effect and limited CNR improvement can be solved.
Owner:NANCHANG HANGKONG UNIVERSITY

Learning abroad application intelligent recommendation method and system based on artificial intelligence

The invention provides an overseas study application intelligent recommendation method and system based on artificial intelligence, and relates to the technical field of overseas study application intelligent recommendation, and the method comprises the steps: obtaining the basic information of an applicant to construct a feature matrix, extracting college admission features through a variational encoder-decoder and an adversarial discriminator; and constructing a dynamic heterogeneous graph structure to execute graph convolution operation to update the matching probability, and analyzing a group decision based on an information cascade effect to obtain an optimal application combination sequence. According to the method, personalized accurate recommendation can be realized, the study abroad application success rate is improved, and the application decision complexity is reduced.
Owner:SHANGHAI TIANQU YUNQI EDUCATION TECHNOLOGY CO LTD

Concatenated code encoding method, concatenated code decoding method, and communication apparatus

This application provides example concatenated code encoding and decoding methods. In one example method, polar code encoding is performed on a to-be-encoded message bit sequence based on a frozen bit set and a message bit set that are of a first polar code, to obtain a first encoded code word whose length is N_o. The first encoded code word is interleaved to obtain a second encoded code word, and optimized low-density parity-check (LDPC) code encoding is performed on the second encoded code word to obtain a concatenated code word. Interleaving includes both outer interleaving and inner interleaving. The concatenated code word is then outputted.
Owner:HUAWEI TECH CO LTD

System and Method for Interactive Robot Action Replanning Using Large Language Models

A robotic controller for controlling a robot according to a sequence of robotic actions. comprises an input interface to receive multimodal inputs specifying instructions for performing a task in audio, video, and a text modality. The controller transforms the multimodal instructions into encodings using a large language model (LLM) encoder and decodes the encodings into a first sequence of robotic instructions and a robot action description of the actions using an LLM decoder. Human feedback input is received corresponding to at least one action in the first sequence of actions and the controller encodes the feedback input with the robot action description. The controller feeds the encoded data along with multimodal features generated from the encodings into the LLM decoder to generate a corrected sequence of actions. The controller is configured to control a robot according to the corrected sequence of actions.
Owner:MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC

Method and apparatus for encoding / decoding image and recording medium for storing bitstream

An image encoding / decoding method and apparatus, a recording medium storing a bitstream and a transmission method are provided. The image decoding method may comprise generating a derived intra prediction mode by performing decoder side intra mode derivation (DIMD) on a current block and storing the derived intra prediction mode. The current block is in a matrix based intra prediction (MIP) mode.
Owner:HYUNDAI MOTOR CO LTD +1

Method and apparatus of encoding / decoding point cloud geometry data sensed by at least one sensor

A method and apparatus for encoding / decoding a point cloud may use any type of sensor following a sensing path. The method obtains coarse representations of sensed points and encodes the sensing path and the coarse representations. The sensing path and coarse representations of points are decoded, and points of the point cloud are reconstructed from the decoded sensing path and the decoded coarse representations.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Intra block copy-based encoding / decoding method, device, and bitstream storage medium

An image encoding / decoding method and apparatus are disclosed. The image decoding method includes acquiring prediction mode information of a current block from a bitstream, decoding an intra block copy prediction mode of the current block using the prediction mode information of the current block, and reconstructing the current block based on the intra block copy prediction mode. The intra block copy prediction mode is at least one of a block copy based SKIP mode, a block copy based MERGE mode or a block copy based AMVP mode.
Owner:ELECTRONICS & TELECOMM RES INST

Unconditional multivariate time series data generation method and device based on diffusion model

The invention relates to the technical field of deep learning, in particular to an unconditional multivariate time series data generation method and device based on a diffusion model, and the method comprises the steps: building a denoising model based on a generation normal form of a diffusion model DDPM through employing a Transform coding and decoding structure, directly predicting original sample data through employing the denoising model in a training process, and carrying out the prediction of the original sample data, the prediction data of the denoising model is associated with the original sample data, so that the trained denoising model focuses on the quality of the prediction data generated by optimization, the generated data does not need to be indirectly recovered through prediction noise, and the quality of the generated multivariate time series data is effectively improved.
Owner:JIANGNAN UNIV