Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1737 results about "Coding decoding" patented technology

Video semantic segmentation method based on time sequence cross attention mechanism

The invention discloses a video semantic segmentation method based on a time sequence cross attention mechanism, and belongs to the field of computer vision and the field of material detection.The video semantic segmentation method comprises the steps that firstly, a video used for training is preprocessed, a frame sequence is extracted, and then a coding-decoding network for multi-level feature extraction and fusion is constructed; according to the method, feature extraction is enhanced through a time sequence cross attention module, network parameters are optimized through weighted IoU loss and binary cross entropy BCE loss, then frame-by-frame prediction segmentation is carried out on a target video by using a trained model, and a multi-classification segmentation result is exported. According to the method, a time sequence cross attention mechanism is integrated into the SAMUNet network, the segmentation precision is effectively improved for image data with time sequences, the time cost and the labor cost of material video processing are greatly reduced, the method can be widely applied to the field of industrial detection, and the product quality and the production efficiency are improved.
Owner:ZHEJIANG UNIV

Edge-deployed semi-supervised anomaly detection method and system for railway track foreign object

Disclosed in the present invention are an edge-deployed semi-supervised anomaly detection method and system for a railway track foreign object. The method comprises the following steps: an edge device encoding and decoding a video stream captured by a camera to obtain an image frame sequence, and performing frame extraction; and using a semantic segmentation model to perform image segmentation on a certain image frame obtained by means of frame extraction, to obtain a railway track region segmentation image. The use of a single image as input may generate an expert model result having a high weight value; however, the determination based on a single image is not stable, multiple consecutive images of the task scene need to be inputted, the frequency of each expert model obtaining the highest weight is computed, and the expert model corresponding to the highest frequency is the final solution. The present invention supports scene-adaptive foreign object detection algorithm automatic selection, and a user can perform selection on the basis of prior knowledge, or selection may be performed by a scene-adaptive automatic algorithm selection method; the user only needs to provide a batch of image data of the current scene, and the optimal algorithm selection can be evaluated.
Owner:GUANGZHOU EMBEDDED MACHINE TECH CO LTD

Method, device, and recording medium for image encoding / decoding

Disclosed herein are a method, an apparatus and a storage medium for image encoding / decoding. In typical image encoding / decoding methods, a decoder-side motion information derivation method may be limitedly used. Therefore, the improvement of encoding efficiency attributable to the decoder-side motion information derivation method may also be limited. In embodiments, a motion information search method used in an inter-prediction mode, an intra block copy mode and an intra template matching prediction mode is disclosed. With the use of various motion search methods, encoding efficiency may be improved.
Owner:ELECTRONICS & TELECOMM RES INST

Method, device, and recording medium for image encoding / decoding

Disclosed herein are a method, an apparatus and a storage medium for image encoding / decoding. Methods for performing more accurate prediction are disclosed when performing prediction using a block vector. A block vector for a target block is derived by template matching, and a prediction block for the target block is determined based on the derived block vector. Various and detailed embodiments for a plurality of pieces of information used to derive a block vector, such as a search area, a block vector candidate list, a reference region, a buffer, and a tree type, are provided.
Owner:ELECTRONICS & TELECOMM RES INST

System and method for robot planning using large language models

A robotic controller for controlling a robot according to a sequence of robotic actions. comprises an input interface configured to receive a plurality of multimodal inputs each specifying instructions for performing a task in a different modality including audio, video, and a text modality. The controller also comprises a multimodal large language model, an action sequence decoder, and a controller. The multimodal LLM includes a multimodal LLM encoder and an LLM decoder. The multimodal LLM encoder is trained with machine learning to transform the multimodal instructions into encodings and the LLM decoder is configured to decode the encodings into a sequence of robotic instructions. The action sequence decoder is trained with machine learning to transform the sequence of robotic instructions into a sequence of actions using a library of robotic skills. The controller is configured to control a robot according to the sequence of actions.
Owner:MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC

Image encoding / decoding method, device, and recording medium for storing bitstream

This disclosure provides an image encoding / decoding method and device and recording medium storing a bitstream. The image decoding method may comprise: constructing a candidate list related to prediction of a current block; generating a prediction block of the current block based on a candidate selected from the candidate list; and generating a reconstructed block of the current block based on the prediction block, wherein the constructing of the candidate list includes adding candidates to the candidate list in an order of a first candidate serving as a spatially neighboring block, a second candidate serving as a temporally neighboring block, and a third candidate serving as a non-neighboring block.
Owner:ELECTRONICS & TELECOMM RES INST

Lightning approaching prediction method and device based on multi-source meteorological data

The invention discloses a thunder approaching prediction method and device based on multi-source meteorological data, and particularly relates to the technical field of thunder disaster prediction.The method comprises the steps that standardization processing of temporal-spatial resolution unification is conducted on meteorological satellite data, radar data and lightning positioning data, and standardized temporal-spatial input data is generated; then spatial-temporal features are extracted through a depth separable 3D convolution module, and a compressed spatial-temporal feature graph is generated; then, a channel-space double attention mechanism (CPCA) is applied to the compressed feature map for feature optimization; and finally, processing the optimized feature map through a coding-decoding structure, and outputting a thunder and lightning probability distribution map. According to the method, the problem of spatial-temporal resolution difference and physical feature mismatching in multi-source data fusion is innovatively solved, the calculation efficiency is remarkably improved through a lightweight network architecture, the feature expression ability is enhanced by using an attention mechanism, and high-precision prediction of sudden thunder and lightning events is realized.
Owner:CHENGDU UNIV OF INFORMATION TECH

Image defogging method based on dynamic wavelet prior and double-domain learning

The invention belongs to the technical field of image processing and deep learning, and particularly relates to an image defogging method based on dynamic wavelet prior and double-domain learning. Aiming at the requirements of all-weather clear imaging in the fields of intelligent traffic systems, safety monitoring and the like, and in order to overcome the defect that a static convolution kernel adopted by a traditional defogging method is difficult to adapt to different haze degradation, the invention provides a method for dynamically generating a convolution kernel by using haze priori contained in a multi-scale wavelet LL sub-band; and an efficient, robust and accurate image defogging model is constructed. According to the invention, based on a multi-scale U-shaped coding-decoding architecture, a dynamic wavelet depth separable convolution module DyWConv is embedded in front of each level of a coder to realize content adaptive feature extraction, and a double-domain feature learning module SPAFormer Block cooperatively utilizing Fourier domain global modulation and wavelet domain multi-scale decomposition is designed. And double-domain features are fully fused through an adaptive gating fusion mechanism, and finally a clear image is reconstructed and output step by step. According to the method, a method for explicitly encoding frequency domain degradation prior into dynamic convolution kernel parameters is innovatively provided, the complementary advantages of Fourier transform and wavelet transform are cooperatively utilized, spatial non-uniform haze can be effectively removed, image details can be recovered, leading performance is achieved in a synthetic data set and a real scene, and the method has a wide application prospect.
Owner:NANKAI UNIV

Multiplying active code decoding method based on SO-ORBGRAND decoder

The invention belongs to the technical field of communication, and particularly relates to a multiplication active code decoding method based on an SO-ORBGRAND decoder. The SO-ORBGRAND decoder outputs soft information by calculating a correct posterior probability of a conjecture code word sequence, and high-performance decoding is realized through soft information exchange and optimization processing between the row decoder and the column decoder. In the iteration process, a method for judging the validity of the code word through the check matrix is designed by utilizing the polarization characteristic of the multiplication active code, so that the iteration is supported to jump out in advance to reduce the complexity. The invention provides a multiplication active code decoding algorithm based on an SO-ORBGRAND decoder, and compared with various non-system multiplication active code decoding algorithms and the performance of traditional long-code polarization codes, the method has remarkable advantages in error correction performance; a method for judging the validity of the code word through the check matrix by utilizing the polarization characteristic design of the multiplication positive code is provided, the complexity is reduced while the judgment success rate is improved, and the implementation overhead is reduced.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Encoding and Decoding Method and Apparatus

An encoding method includes first obtaining a to-be-encoded picture block, then determining encoding complexity of the to-be-encoded picture block based on a plurality of prediction modes, next determining a quantization parameter of the to-be-encoded picture block based on the encoding complexity, and after that encoding the to-be-encoded picture block based on the quantization parameter to generate a bitstream. The plurality of prediction modes include at least one of block-level prediction, point-level prediction, and intra block copy prediction, and the encoding complexity indicates encoding difficulty of the to-be-encoded picture block.
Owner:HUAWEI TECH CO LTD

Method and apparatus for encoding / decoding image

To provide a video encoding / decoding method for performing intra-prediction on the basis of availability of a reference pixel so as to improve encoding performance.SOLUTION: An intra-prediction method for a color copy mode checks a reference pixel region designated to obtain correlation information, determines a reference pixel processing setting based on a determination of availability of the reference pixel region, and performs intra-prediction according to the determined reference pixel processing. In addition, a candidate list for motion information prediction of the current block is generated, a control point vector of the current block is derived based on the candidate list and a candidate index, a motion vector of the current block is derived based on the control point vector of the current block, and inter prediction is performed on the current block.SELECTED DRAWING: Figure 14
Owner:INST OF IMAGE TECH INC

Efficient deep supervised distillation method and device applied to medical image segmentation

The invention relates to the technical field of medical image segmentation, in particular to a high-efficiency deep supervised distillation method and device applied to medical image segmentation, and the device comprises a pre-training coding module, a high-efficiency multi-scale fusion block, a multi-scale grouping attention gate, a high-efficiency frequency domain boundary extraction block, a high-efficiency fusion up-sampling block, and knowledge alignment multi-scale distillation. According to the invention, while the multi-scale features are effectively captured and fused, the feature expression capability in jump connection is enhanced; besides, the boundary prediction capability of the model is improved by using frequency domain information, and the multi-scale knowledge of the teacher model and the student model is gradually aligned in the encoding-decoding stage, so that the lightweight student model can still keep excellent segmentation performance under the condition of greatly reducing the calculation cost, and the method has extremely high compatibility, and is suitable for popularization and application. The method can be adaptive to multi-layer visual encoders of different scales, can be adaptive to output of different layers of the encoders, and effectively reduces the calculation amount and the parameter amount while guaranteeing the high segmentation performance.
Owner:DONGGUAN UNIV OF TECH

Graph-based unsupervised fraud detection method

The invention relates to a graph-based unsupervised fraud detection method. The method comprises the following steps: acquiring original graph data; community division is carried out on nodes of the topological structure of the preprocessed graph data; generating a super node graph; a reconstructor network of a coding-decoding structure is trained, and a reconstructor model is obtained through iterative training; calculating a reconstruction score of the super node; sampling a node 1-ego network; neighbor nodes in the 1-ego network are divided into a same-proton graph and a different-proton graph; training a graph neural network; iteratively training to obtain a graph neural network model; obtaining an isoproton diagram code and an isoproton diagram code; generating a second abnormal score of the target node; and marking the nodes with the comprehensive abnormal scores exceeding an abnormal threshold value as fraud nodes. Compared with the prior art, the method has the advantages of detecting aggregation type and sparse type fraud at the same time and the like.
Owner:SHANGHAI UNIV

Feature encoding / decoding method and apparatus, and recording medium storing bitstream

Provided are a feature encoding / decoding method and apparatus, and a computer-readable recording medium generated by the feature encoding method. The feature decoding method may comprise obtaining first information on a maximum number of internal layers allowed in a coded feature sequence (CFS) from a bitstream, obtaining second information on an identifier for each internal layer and third information on the number of channel layers of the internal layer from the bitstream, based on the first information, obtaining fourth information about an identifier for each channel layer from the bitstream based on the third information, and reconstructing the channel layer based on the fourth information.
Owner:LG ELECTRONICS INC

Video acquisition method and device based on modular low coupling

The invention discloses a video acquisition method and device based on modular low coupling, and relates to the technical field of video processing. The method comprises the following steps: initializing a video processing system, creating a unified system resource management interface, and realizing centralized management of video acquisition, encoding, decoding, OSD control and NPU push modules; based on a modular design principle, video acquisition, coding, decoding, OSD control and NPU push functions are packaged into independent modules, and data interaction between the modules is realized through a standardized interface. The video processing system is modularly designed, and a unified standardized interface is provided, so that the effect of greatly reducing project code coupling is achieved, a developer does not need to pay attention to underlying logic when calling the interface, and module function independence is guaranteed; by means of the modular interface design, developers can rapidly integrate platform functions, development work can be completed without deeply knowing the whole code library, and therefore the project development efficiency is improved.
Owner:XIAN XINCHEN ELECTRONIC TECH CO LTD

Sea-land island remote sensing image segmentation method based on feature difference enhancement fusion

The invention discloses a sea-land island remote sensing image segmentation method based on feature difference enhancement fusion, and the method comprises the steps: collecting sea-land island remote sensing image data, and constructing a sea-land island remote sensing image data set; constructing an initial sea-land island remote sensing image segmentation model; calculating a loss function value of the model, performing back propagation, and training to obtain a sea-land island remote sensing image segmentation model; and outputting a segmentation result graph based on the trained model. According to the method, a coding-decoding architecture is adopted, a difference attention fusion module is introduced in a coding stage, and the two features are fused based on complementarity of global and local features so as to enhance the expression ability of the model; a cross fusion module is introduced in the decoding stage, multi-scale feature fusion is performed on features from different coding stages, and the segmentation precision is improved; and finally, outputting a segmentation prediction result through a full connection layer, and realizing accurate training optimization in combination with a multi-term auxiliary loss function, thereby improving the efficiency and accuracy of sea-land island remote sensing image segmentation.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

Image text description generation method, electronic equipment and readable storage medium

The invention provides an image text description generation method, electronic equipment and a readable storage medium. According to the method, the object perception prototype learning module and the global context feature extraction module are introduced, so that fine-grained information and global semantic understanding in the image are effectively balanced. The visual backbone network module can extract multi-scale and multi-level image features and perform fusion, thereby enhancing the expression ability of the image features. The object perception prototype learning module further extracts an object prototype from the fusion features to ensure that the model can accurately capture key objects and attributes thereof in the image, and the global context feature extraction module ensures that the overall context of the image is fully understood. On the basis, the encoding and decoding module combines the global context and the object prototype to generate the text description, so that the semantic splitting phenomenon in the traditional method is avoided, and the detail information in the image is effectively reserved, thereby improving the accuracy and integrity of the image description.
Owner:WUHAN UNIV

Image encoding / decoding method and apparatus

An image encoding / decoding method and apparatus according to the present invention may: determine an intra prediction mode of a current block; determine one or more reference lines of the current block from among a plurality of reference line candidates available for the current block; generate a prediction sample of the current block on the basis of the intra prediction mode and the one or more reference lines, which have been determined; and correct the generated prediction sample.
Owner:KT CORP

Coagulant addition prediction method based on fusion of sparse coding and graph space-time attention

The invention discloses a sparse coding and graph space-time attention fused coagulant addition prediction method, and relates to the technical field of time sequence data modeling prediction. According to the method, through the combination of a Transform encoding-decoding architecture and a sparse self-attention and top-k sparse strategy, the missing data complementation precision is effectively improved, and high-quality and missing-free input data is provided for subsequent prediction modeling; in the aspect of spatial feature expression, node association of multiple subsystems of a water production system is captured by means of a graph attention network, and by dynamically calculating association weights among nodes and aggregating neighbor features, the information entropy of spatial features is improved, and the expression ability of the spatial features to a multi-node coupling relationship is greatly enhanced; in the aspect of time sequence prediction performance, a time attention mechanism can dynamically focus on key time sequence nodes, the utilization rate of key time sequence features is further improved, and the accuracy of coagulant adding prediction is improved.
Owner:CHENGDU QIANJIA TECH CO LTD

A method, an apparatus and a computer program product for video encoding and video decoding

The embodiments relate to a method for encoding / decoding. The encoding method (900) comprises receiving a video sequence (1505) comprising a first frame and a second frame; encoding (1510) the first frame into a first coded frame using a first coding method (901); reconstructing (1515) a first decoded frame corresponding to the first coded frame; deriving (1520) one or more optimizing parameters to adjust a traditional filter (1110), wherein the optimizing parameters reduce distortion of the first decoded frame to produce a first filtered frame; filtering (1525) the first decoded frame with the traditional filter (1010, 1110); encoding (1530) the second frame into a second coded frame by a second set of algorithms of the second coding method (906, 1210) and by using the first filtered frame directly or indirectly for prediction; and signalling (1535) said one or more optimizing parameters. The embodiments also relate to apparatuses for encoding / decoding.
Owner:NOKIA TECHNOLOGIES OY

Finite geometric code decoding method and system based on generalized check matrix

PendingCN120880462AAlgebraic geometric codesPhase-modulated carrier systemsTheoretical computer scienceBpsk modulation
The invention discloses a finite geometric code decoding method and system based on a generalized check matrix, and relates to the communication technology, and the method comprises the following steps: constructing a corresponding generalized check matrix according to an algebraic structure of a multi-step large number logic decodable finite geometric code; sending a finite geometric code sequence, transmitting the sequence after coding and BPSK modulation, and receiving a posterior probability log-likelihood ratio sequence by a receiving end; and carrying out iterative decoding on the finite geometric code received by the receiving end based on the constructed generalized check matrix. According to the invention, a new generalized check matrix is constructed according to the algebraic structure of the finite geometric code, and multi-step large-number logic iterative decoding is carried out on the finite geometric code based on the matrix, so that the iterative decoding performance of the multi-step large-number logic decodable finite geometric code is improved.
Owner:THE 20TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORP

Image encoding / decoding method and apparatus, and recording medium storing bitstream

This disclosure provides an image encoding / decoding method and device and recording medium storing a bitstream. The image decoding method may comprise: constructing a candidate list related to prediction of a current block; generating a prediction block of the current block based on a candidate selected from the candidate list; and generating a reconstructed block of the current block based on the prediction block, wherein the constructing of the candidate list includes: generating a candidate list for each of two or more types classified based on at least one coding parameter; and generating an integrated candidate list by adding at least one candidate selected in the candidate list for each type in a predetermined order, and wherein the at least one coding parameter is a motion information derivation method, and the two or more types include a first type of a spatially neighboring block, a second type of a temporal neighboring block and a third type of a non-neighboring block.
Owner:ELECTRONICS & TELECOMM RES INST

Skin LC-OCT image segmentation method based on improved GAN and multi-feature fusion

The invention discloses a skin LC-OCT image segmentation method based on improved GAN and multi-feature fusion. The skin LC-OCT image segmentation method comprises the following steps: collecting a skin B-scan sample by adopting LC-OCT equipment and storing the skin B-scan sample as an image; the method comprises the following steps: constructing a generator fusing reflection filling, multi-scale coding-decoding, self-attention and residual blocks, cooperating with a multi-scale fusion discriminator and a weight sharing twin network, and realizing high-fidelity denoising and enhancement through joint optimization of perception loss, multi-scale SSIM loss and Wasserstein adversarial loss; preprocessing and positioning a skin area based on a threshold value or morphology and normalizing the skin area; gradient is extracted from the de-noised image, and dynamic path search is carried out in combination with a gradient difference minimum criterion and a random forest; and after adaptive sampling and spline fitting are carried out on the path, candidates are generated through a random forest, and fracture or dislocation of a focus area is corrected. Therefore, accurate segmentation of the LC-OCT image of the skin layer can be realized in a high-noise environment, a high-quality enhanced data set can be constructed, and a reliable basis is provided for early diagnosis of skin diseases.
Owner:KERNEL MEDICAL EQUIP CO LTD

Joint multi-modal entity relationship extraction and generation method based on multi-view comparative learning

The invention discloses a combined multi-modal entity relationship extraction and generation method based on multi-view comparative learning, and particularly relates to the technical field of entity relationship extraction. The method comprises the following steps: converting triples of entity relationships in all extracted texts into a sequence consisting of position indexes of a head entity and an entity type thereof, a tail entity and an entity type thereof and a relationship between two entities, and generating a target index sequence from end to end in multi-modal input through a BART-based coding-decoding model; three positive samples are constructed for each training sample based on entity, image and context enhancement, a multi-view comparative learning algorithm is introduced, the algorithm adopts a cross entropy target of in-batch negative samples to minimize the distance between the positive samples, and intervals of a head entity and a tail entity in a sentence are specified through a target index sequence to obtain a multi-view comparative learning algorithm; and a category of the relationship so that the multi-modal representation can capture semantic similarities between samples with similar entities and relationship mentions.
Owner:NANJING UNIV OF SCI & TECH

Learning-based point cloud geometry compression framework

In one implementation, geometry of a point cloud is encoded / decoded. On the encoder side, the encoder determines a first feature representing a voxel occupancy status of a current level and / or one or more finer levels of the point cloud, based on the voxel occupancy status of the current level and / or the finer levels; determines a second feature representing prediction of the voxel occupancy status of the current level and / or the finer levels of the point cloud; determines a third feature associated with the voxel occupancy status of the current level and / or the finer levels, based on the first feature and the second feature; and encodes the third feature. On the decoder side, the first feature is decoded from a bitstream, the second feature is determined similarly as the encoder side, and the third feature is determined based on the first feature and the second feature.
Owner:INTERDIGITAL VC HOLDINGS INC

Method and apparatus for encoding / decoding image and recording medium for storing bitstream

An image encoding / decoding method and apparatus, a recording medium for storing a bitstream and a transmission method are provided. The image decoding method comprises generating a chroma mode list of a current chroma block, deriving a chroma intra prediction mode of the current chroma block based on the chroma mode list, and generating a prediction block of the current chroma block based on the chroma intra prediction mode. The chroma mode list may comprise at least one of a default mode, a derivation based chroma mode or a direct mode.
Owner:HYUNDAI MOTOR CO LTD +1