Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

174 results about "Spatial encoding" patented technology

Spatial Encoding. Spatial encoding is probably the most well known and the most intuitive coding method. When spatially encoding, the power (amplitude) of a sample at a particular point in time or space is recorded. ie. over time for audio waves and over space for images.

Underwater image enhancement method based on dual-path feature decoupling and gating fusion

The invention discloses an underwater image enhancement method based on dual-path feature decoupling and gating fusion. The method comprises the following steps: firstly, constructing a training data set; then, an encoder-bottleneck layer-decoder is used as a trunk, and a dual-path encoding and decoding mechanism is adopted to construct an underwater image enhancement model; the encoder is used for extracting direction encoding features and space encoding features; the bottleneck layer is used for extracting global and local features and fusing and outputting bottleneck features; the decoder is used for decoding and multi-scale refinement so as to reconstruct and obtain an underwater enhanced image; then training is carried out to obtain a trained underwater image enhancement model; and finally, the device is deployed to acquire the underwater image in real time for underwater image enhancement. According to the method, independent modeling is carried out for anisotropic scattering / edge attenuation and spatial non-uniform atomization / local brightness imbalance, and mutual interference of different degradation causes in the same feature space can be reduced.
Owner:SOUTH CHINA AGRICULTURAL UNIVERSITY

End-to-end polarization hyperspectral image classification method and system

The invention discloses an end-to-end polarization hyperspectral image classification method and system, belongs to the technical field of deep learning and optical imaging, and solves the technical problems of low reconstruction process speed, limited precision, low information utilization rate in a classification process and weak feature extraction capability in the prior art. The method comprises the following steps: carrying out target shooting based on a snapshot type space coding hyperspectral polarization imaging system, and carrying out system coding on a shot image to obtain two-dimensional aliasing data; reconstructing the two-dimensional aliasing data into a polarization hyperspectral data cube by using the trained reconstruction network; training the classification network based on the polarization hyperspectral data cube to obtain a trained classification network; and performing joint fine tuning on the trained reconstruction network and the trained classification network to obtain a polarization hyperspectral image classification model, and performing polarization hyperspectral image classification by using the polarization hyperspectral image classification model. The method is used for realizing high-quality reconstruction and accurate classification of the polarization hyperspectral image.
Owner:JILIN HAIYUNTIAN ZHIHUI TECHNOLOGY CO LTD

Three-dimensional radar echo reflectivity variational auto-encoder pre-training method

The invention discloses a three-dimensional radar echo reflectivity variational auto-encoder pre-training method, which belongs to the technical field of meteorological radar data analysis, and comprises the following steps: extracting advanced features from three-dimensional high-resolution radar data through a space encoder based on an attention mechanism; mapping the advanced features into determined probability distribution parameters using a probabilistic encoder; sampling according to the distribution parameters by using a re-parameterization technique to obtain continuous potential variables; mapping the potential variables back to a pixel space by using a probability decoder based on an attention mechanism to complete data reconstruction; and finally, constructing a mixed loss function consisting of a mean square error and KL divergence by utilizing a reconstruction result and a distribution parameter to train the model. According to the invention, by optimizing the network architecture and introducing the attention mechanism, high-quality reconstruction of high-resolution three-dimensional radar data is realized while extremely low video memory requirements and low parameter quantity are ensured, and the practical value of data synthesis and enhancement is remarkably improved.
Owner:CHENGDU UNIV OF INFORMATION TECH +2

Lightweight SDN attack detection method based on multi-scale iterative attention

The invention discloses a lightweight SDN (Software Defined Network) attack detection method based on multi-scale iterative attention, relates to the technical field of network security, and solves the problem that an SDN attack detection method based on deep learning in the prior art is insufficient in feature selection static state and spatial modeling and gives consideration to both lightweight and high precision. The method is based on a feature contribution degree evaluation mechanism, the most critical features for attack discrimination are screened out in real time, redundant information is eliminated, and the calculation burden is reduced. Moreover, the attack feature map is generated through normalization, time window overlapping slicing and multi-channel space coding, so that the perception capability of a complex attack mode is improved. Besides, a multi-scale iteration attention mechanism is embedded in a lightweight network architecture, key features are highlighted and redundant information is suppressed through multi-granularity convolution extraction and iteration weight fusion, and both lightweight and high-precision detection are realized, so that the method is suitable for real-time network environment and edge device deployment.
Owner:ELECTRIC POWER RES INST OF GUANGXI POWER GRID CO LTD

Single-pixel calculation hyperspectral imaging system and method

The invention provides a single-pixel calculation hyperspectral imaging system and method, and the technical scheme of the invention comprises a spatial light modulator which is used for carrying out the spatial coding of a light beam from an imaging scene; and the reconfigurable spectrum detector is used for detecting the light beams after space coding. The reconfigurable spectrum detector is provided with a group of spectral response functions which are tunable along with external bias voltage, broadband and correlative, and the reconfigurable spectrum detector forms a core sensing unit of the micro-computing spectrometer. A barrel signal sequence is obtained by synchronously collecting light intensity integral values under different space coding patterns and different bias voltages. And finally, in a data processing module, a spectral data cube of the target scene is jointly solved from the bucket signal sequence according to the spatial coding matrix and the spectral response function set by utilizing a calculation reconstruction network. According to the invention, two computational imaging technologies are creatively fused, the structure is very compact, and real-time and wide-spectrum spectral imaging can be realized at a low sampling rate.
Owner:LIAONING UNIVERSITY

Deep learning prediction method and system for multi-source space-time lattice point data

The invention relates to the technical field of deep learning, and discloses a deep learning prediction method and system for multi-source space-time lattice point data, and the method comprises the steps: obtaining live lattice point data, forecast lattice point data and static lattice point data, and carrying out the preprocessing, thereby obtaining uniform lattice point feature data; and generating missing mask lattice point data, generating confidence coefficient lattice point data according to observation coverage information, interpolation distance information, a time-space consistency test result and a forecast aging attenuation rule, and forming model input lattice point data. And constructing a training sample and generating forecast aging coding data. And constructing and training a deep learning prediction model, wherein the model comprises a space coding network, a time coding network and a forecast aging segmentation prediction network. In the training stage, confidence coefficient is used for weighting regression loss. In the reasoning stage, a release prediction result is obtained. According to the method, the learning stability of the multi-source grid point data under the missing measurement and filling conditions is improved, and the reliability of continuous value prediction and grading alarm probability under different prediction time periods is improved.
Owner:贵州省气象台

Generative reverse face recognition method based on text guidance

The invention provides a text guidance-based generative reverse face recognition method. The method comprises the following steps of: obtaining an initial latent vector and a fine tuning generator by using a pre-trained editing encoder and an original face image; taking the initial latent vector as an initial vector during gradient updating, and starting to circularly update until a final latent space code is obtained; and inputting the subsurface space code into the fine-tuned generator, and taking the generated image as a protected image corresponding to the original face image. The total loss function comprises an adversarial loss for explicitly promoting diversity and a loss function of a visual effect; the fine-tuned image generated by the generator is closer to an auxiliary image randomly selected from the auxiliary data set so as to realize protection of dynamic tracking of face recognition; meanwhile, it is ensured that the image generated by the fine-adjusted generator follows the specification of target text prompt and keeps the visual similarity with the original face image. The method has excellent performance in the aspect of preventing the face image of the user from being identified by the dynamic FR strategy.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Edge cloud computing resource allocation optimization method based on deep learning

The invention relates to the field of intelligent scheduling allocation, in particular to an edge cloud computing resource allocation optimization method based on deep learning, which adopts a space-time prediction algorithm based on multi-head attention and gating mechanism optimization to design time coding and space coding. The spatial relationship and interaction between time sequence characteristics of the computing power load and edge server nodes are captured, and meanwhile, a multi-head attention mechanism and expansion causal convolution are combined, so that instantaneous computing power load fluctuation can be captured, and the long-term trend of the computing power load can be mined; therefore, a reliable basis is provided for subsequent computing power scheduling by predicting an accurate computing power load. The invention designs an alternating direction multiplier method based on genetic algorithm optimization, which is not only suitable for a nonlinear and multi-constraint optimization problem, but also can be expanded to a larger-scale distributed edge node cloud computing system, and meanwhile, a global optimal solution is quickly approached through the genetic algorithm, so that the quality of an initial solution is improved, and model convergence is accelerated; and the distributed collaborative allocation scheduling efficiency is improved.
Owner:MIANYANG TEACHERS COLLEGE

Multi-temporal remote sensing crop extraction method in combination with MoCo self-supervised learning

The invention discloses a multi-temporal remote sensing crop extraction method in combination with MoCo self-supervised learning. Relates to the technical field of agricultural resource monitoring, in particular to a multi-temporal remote sensing crop extraction method combined with MoCo self-supervised learning. According to the method, label-free satellite images are utilized, MoCo self-supervised learning is adopted to carry out label-free pre-training on a space encoder, so that a crop classification model can still learn spatio-temporal characteristic representation with discriminative power under limited label data, and the classification precision problem caused by insufficient labeled samples is effectively relieved. The method comprises the following steps: acquiring a multi-temporal crop label-free remote sensing satellite image data set and a label data set; constructing a time-phase crop classification model: pre-training a spatial feature encoder by MoCo self-supervised learning; performing supervised learning on the crop classification model by adopting the labeled data set to obtain a final time phase crop classification model; and classifying the multi-temporal crops through the final temporal crop classification model.
Owner:JILIN AGRICULTURAL UNIV

A semi-implicit neural map construction method based on grid-like cell spatial coding

The application discloses a semi-implicit neural map construction method based on grid-like cell space coding. The method uses grid-like cell space coding to abstractly encode a three-dimensional space, inputs the coding result into a neural network for decoding, and generates a new visual view through rendering. The grid-like cell space coding on the three-dimensional space can improve the correlation between data, eliminate redundant information in the neural map, realize compression sensing of the environment, maximize the use of information, and thus improve the map reconstruction quality. The semi-implicit neural map construction method can be applied to technical fields such as medical imaging, automatic driving, game development or indoor design, and can automatically generate a three-dimensional scene model according to an input two-dimensional image.
Owner:HANGZHOU DIANZI UNIV

A real scene video deblurring system and method based on a single-step video diffusion model

The application discloses a real scene video deblurring system and method based on a single-step video diffusion model, comprising an encoding module, a denoising module and a decoding module, wherein the encoding module is used for respectively performing latent space encoding on each frame in a blurred video sequence to be recovered, generating a frame-by-frame latent space representation corresponding to the input video frame by frame; the denoising module is used for performing single-step denoising on the frame-by-frame latent space representation to obtain a latent space representation corresponding to a clear video; the decoding module is used for decoding the latent space representation corresponding to the clear video into image frames frame by frame and outputting according to the original time sequence of the input video to obtain a deblurred video result. Through frame-by-frame latent space encoding, frame-by-frame blur differences can be preserved, and through single-step diffusion distillation, reasoning delay can be reduced, and deblurring quality and reasoning efficiency are considered.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

An end-to-end polarization hyperspectral image classification method and system

An end-to-end polarization hyperspectral image classification method and system, belonging to the field of deep learning and optical imaging technology, solves the technical problems of slow reconstruction speed, limited accuracy, low information utilization, and weak feature extraction capabilities in existing technologies. The method involves capturing images of a target using a snapshot-style spatially coded hyperspectral polarization imaging system. The captured images are then encoded by the system to obtain two-dimensional aliased data. A trained reconstruction network is used to reconstruct the two-dimensional aliased data into a polarization hyperspectral data cube. The classification network is then trained based on the polarization hyperspectral data cube to obtain a trained classification network. The trained reconstruction network and the trained classification network are jointly fine-tuned to obtain a polarization hyperspectral image classification model, which is then used to classify polarization hyperspectral images. This invention achieves high-quality reconstruction and accurate classification of polarization hyperspectral images.
Owner:JILIN HAIYUNTIAN ZHIHUI TECHNOLOGY CO LTD

X-ray communication device and method based on laser-driven matrix photocathode

ActiveCN121601521BX-ray tube electrodesPhoto-emissive cathodes manufactureTarget arrayPhotocathode
The application provides an X-ray communication device and method based on a laser-driven matrix photocathode. The device is applied to the technical field of X-ray tubes and comprises a laser control circuit, a matrix transmission photocathode array, a collimating microchannel plate, a matrix transmission anode target array, a vacuum tube shell and a beryllium window. The method comprises: generating a dynamic pattern, loading data information to be input into an optical signal; and exciting photoelectron emission with spatial resolution characteristics at a corresponding photocathode pixel area; each channel cluster is aligned with a single photocathode pixel at the front end; after the electrons enter the microchannel, continuous secondary electron emission occurs under the action of the electric field on the inner wall of the channel, and an avalanche effect is generated; and through mechanisms such as bremsstrahlung radiation, a bundle of microfocus X-rays is generated at each target point. In this way, the technical problems of the X-ray source in the prior art, such as the limitation in spatial coding capability, response speed, integration degree and communication dimension, can be solved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Intelligent lighting data transmission system based on visible light communication

This invention relates to the field of visible light communication technology and discloses a smart lighting data transmission system based on visible light communication. The system includes an optical signal processing module, a transmission management module, and a terminal control module. The optical signal processing module acquires and processes visible light signals, outputting a spatially encoded light intensity signal. The transmission management module receives this signal, decodes it through a signal analysis unit to obtain the original data stream signal and channel interference characteristic signal, processes it through a channel optimization unit to obtain the optimal modulation parameter signal, and outputs a dynamic bandwidth allocation signal through a transmission strategy unit. In the terminal control module, an adaptive modulation execution unit adjusts the light transmitter drive parameters according to the optimal modulation parameter signal, a lighting array control unit adjusts the working state of the lighting units and collects real-time link status signals, and a link monitoring module outputs a communication quality index. This system is suitable for scenarios integrating smart lighting and data transmission.
Owner:YIWU TORCH ELECTRONIC CO LTD

A source-load joint probability prediction method and system of a physically constrained graph attention network

The application discloses a source-load joint probability prediction method and system of a physically constrained graph attention network. The method collects multi-dimensional feature data of source-load nodes in a prediction area to construct an initial node feature matrix. A similarity matrix is generated through differentiable graph structure learning. A dynamic adjacency matrix is generated through normalization and introduction of a sparse mask. Spatial feature aggregation is performed through a multi-head graph attention network to obtain node spatial encoding. The node spatial-temporal hidden state is output through an encoder. The node spatial-temporal hidden state is input into a probability prediction head to output Gaussian distribution parameters of the source-load node power. A joint loss function is constructed. The joint loss function is used for soft constraint training to output a probability prediction result. Posterior projection hard constraint correction is performed in an inference stage to obtain a corrected prediction result. The application solves the problems of lack of physical consistency and inability to quantify uncertainty in the prior art.
Owner:BEIJING NORTH STAR DIGITAL REMOTE SENSING TECH CO LTD +1

Method for establishing entity object relationship of urban governance element data

PendingCN122045322ANatural language data processingGeographical information databasesKeyword analysisSpatial code
The invention discloses a method for establishing an entity object relationship of urban governance elements. The method comprises the following steps of: firstly, dividing urban governance multi-source data into spatial data and non-spatial data; giving spatial codes to the spatial data according to a spatial grid where a central point of the spatial data is located; especially for the planar elements, a tolerance calculation mechanism is introduced, when spatial overlay analysis is carried out, only when the geometric overlapping relation between the planar elements and other spatial elements meets the preset tolerance calculation condition, it is judged that the planar elements have the inclusion relation and correlation is established, and therefore boundary overflow errors caused by surveying and mapping precision are corrected. And for non-spatial data, accurately mounting the non-spatial data to the established spatial entity through address standardization matching of place name address objects or keyword analysis. And finally, on the basis of a unique space code formed by combining the professional identification code, the space identification code and the time identification code, constructing an urban governance element entity object. According to the method, the problems of inaccurate spatial logic judgment and difficult association of non-spatial data in traditional data fusion are effectively solved, and accurate fusion and efficient indexing of urban governance data are realized.
Owner:BEIJING THUPDI PLANNING DESIGN INST

An intelligent commercial site selection method, device and medium based on big data analysis

ActiveCN121998382BData streamData set
The application discloses an intelligent commercial site selection method and device based on big data analysis and a medium, relates to the technical field of big data analysis, and comprises the following steps: collecting a commercial site selection coupling data set, performing feature alignment, and forming a standard space-time data stream; adopting an Apache Flink stream processing engine to perform complex event mode recognition on the standard space-time data stream, and forming a reference point address digital portrait; inputting the reference point address digital portrait and the commercial site selection coupling data set into a deep learning site selection model; performing time series modeling and space encoding through a feature coding layer; performing space clustering analysis and similarity calculation through a similarity measurement layer; and outputting a site selection similarity cloud map. Through the Apache Flink stream processing engine and the deep learning site selection model, the application enhances the fusion precision between multi-source data and significantly improves the dynamic prediction capability of commercial site selection.
Owner:ZHEJIANG KESHU STORE TECHNOLOGY CO LTD

Video target detection method and system based on hybrid Transform-Mama

The invention relates to the technical field of target detection, and provides a video target detection method and system based on hybrid Transform-Mama. The method comprises the following steps: based on all frame images in a video to be detected, generating a Token feature sequence by adopting a shared feature extractor, fusing the Token feature sequence and a position code, and inputting the fused Token feature sequence and the position code into a spatial adaptive deformable Transform encoder to obtain spatial encoder features of all frames; splicing the space encoder features of all the frames and then inputting the spliced space encoder features into a time sequence cascade bidirectional Mama encoder to generate space-time encoder features of all the frames; inputting the space-time encoder features of all frames and the target query into an entangled Mama-Transform decoder, and enriching instance-level context information of the target query through query-feature interaction and fine granularity alignment to obtain space-time decoder features; and inputting the space-time decoder features into a shared feed-forward network for classification and bounding box regression to obtain a target detection result of each frame.
Owner:QINGDAO UNIV OF SCI & TECH

A radio frequency coil design method for transmit array spatial encoding imaging

The application relates to a radio frequency coil design method for transmit array spatial encoding imaging, and belongs to the technical field of magnetic resonance imaging. In view of the problems of low single-dimension coding efficiency, strong coupling interference between coils and nonlinear phase gradient of a traditional TRASE technology, axial and radial coding coils are designed based on a target field method and a magnetic dipole method, two-dimensional radio frequency phase coding is realized through reverse current solenoid decoupling combination, and excitation and signal receiving are realized in combination with a saddle coil. Through optimization of the coil structure and decoupling design, the technical scheme achieves multi-dimensional efficient coding, high-precision phase gradient, strong anti-interference and system light weight, and is suitable for the portability and dynamic monitoring scene of low-field magnetic resonance.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Feedforward event camera three-dimensional reconstruction method and system based on spatiotemporal feature aggregation

The application discloses a kind of feedforward event camera three-dimensional reconstruction method and system based on space-time feature aggregation, comprising: obtaining at least two asynchronous event streams, convert each event stream into space-time voxel tensor, space-time voxel tensor includes multiple time boxes, each time box corresponds to the event accumulation in a short time fixed time slice;Space-time voxel tensor is input into time attention encoder, the feature of each spatial position is aggregated in time sequence on different time boxes by self-attention mechanism, obtain space-time feature map rich in time context information;Space-time feature map is input into the space encoder-decoder based on feedforward architecture, and the globally aligned three-dimensional point graph is generated by regression.The method and system can effectively extract the motion information and geometric clues in event data, can accurately predict three-dimensional point graph, and are significantly better than existing event camera methods in depth estimation, camera pose estimation and three-dimensional reconstruction tasks.
Owner:ZHEJIANG UNIV +1

Control method, device and storage medium of blood pressure detection system

PendingCN122451854ACardiac cycleSimulation
The application discloses a blood pressure detection system control method, equipment and storage medium, and belongs to the technical field of blood pressure detection. The method comprises the following steps: performing multi-modal decomposition on a to-be-detected physiological signal to form a space-time input matrix, inputting the space-time input matrix into a space-time fusion model, calculating the attention weights between channels in the space-time input matrix through a space encoder, performing weighted aggregation on channel features based on the attention weights to obtain a feature sequence, identifying long-range time sequence dependence in the feature sequence through a time sequence encoder to obtain a space-time fusion feature vector, and performing regression mapping on the space-time fusion feature vector through a multi-task output layer to output blood pressure detection data. The application adaptively enhances the signal components related to blood pressure through a space-time encoder, and suppresses noise interference. Meanwhile, based on the multi-task output layer, the robustness to posture changes and breathing interference is improved, and the accuracy of blood pressure detection is improved.
Owner:HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Remote sensing interpretation method and system integrating multi-source spatiotemporal spectral features and visual models

This invention relates to a remote sensing interpretation method and system that integrates multi-source spatiotemporal spectral features and a visual model. First, optical imagery, harmonic data, and synthetic aperture radar (SAR) data of the target area are acquired and preprocessed. Next, spatiotemporal features of the optical imagery and harmonic data are extracted separately and fused using a cross-attention mechanism to generate fused features. Then, the fused features and SAR data are subjected to image serialization, temporal encoding, and spatial encoding processing, and spatiotemporal features are extracted using a self-attention mechanism. Finally, the spatiotemporal features are decoded and the multi-source feature convolution results are fused to generate pixel-level land cover classification results. This invention introduces harmonic data to capture the temporal variation trend of ground features, utilizes the visual Transformer self-attention mechanism to capture spatiotemporal interdependencies, and combines multi-source data to capture multi-dimensional features of ground features, avoiding the limitations of a single data source. Furthermore, the fusion of multi-source temporal features allows the model to adapt to different geographical environments, improving generalization performance.
Owner:SOUTH CHINA NORMAL UNIV

Ultraviolet light communication system and demodulation method based on transverse effect position sensitive detector

The invention discloses an ultraviolet light communication system based on a transverse effect position sensitive detector, and belongs to the technical field of optoelectronic devices. Comprising a space code modulation module, an imaging lens and a signal demodulation module. The spatial coding modulation module is composed of a 2 * 2 high-frequency flicker ultraviolet LED array, a driving circuit and an imaging lens, four LED units are defined as four spatial coding states respectively, information to be transmitted is converted into an 8-bit binary sequence according to an ASCII coding rule, and the 8-bit binary sequence is split into four groups of 2-bit sub-codes; each group of sub-codes realizes information coding based on a two-dimensional space position by controlling the time sequence lightening of a single LED unit; the signal demodulation module comprises a position sensitive detector and a signal processing circuit thereof. The signal processing circuit is configured to restore a binary sub-code corresponding to an incident light spot through a method of comparing amplitude extreme values of four electrode output signals of the position sensitive detector; and then splicing the demodulated sub-codes into a complete data frame. The method is suitable for short-distance ultraviolet light communication in a complex environment.
Owner:BEIJING UNIV OF TECH

Signal reshaping for high dynamic range signals

In a method to improve backwards compatibility when decoding high-dynamic range images coded in a wide color gamut (WCG) space which may not be compatible with legacy color spaces, hue and / or saturation values of images in an image database are computed for both a legacy color space (say, YCbCr-gamma) and a preferred WCG color space (say, IPT-PQ). Based on a cost function, a reshaped color space is computed so that the distance between the hue values in the legacy color space and rotated hue values in the preferred color space is minimized. HDR images are coded in the reshaped color space. Legacy devices can still decode standard dynamic range images assuming they are coded in the legacy color space, while updated devices can use color reshaping information to decode HDR images in the preferred color space at full dynamic range.
Owner:DOLBY LABORATORIES LICENSING CORP

Multi-modal full body pose tracking

Techniques and systems are provided for pose prediction. For instance, a process can include combining image features detected from an obtained image with estimated image features to generate combined features; generating temporally encoded features by temporally encoding the combined features; combining detected motion tracking information with estimated motion tracking information to generate combined motion tracking information; generating temporally encoded motion tracking information by temporally encoding the combined motion tracking information; generating spatially encoded multi-modal information by spatially encoding the temporally encoded features and the temporally encoded motion tracking information; and predicting a body pose by regressing the spatially encoded multi-modal information.
Owner:QUALCOMM INC

Method for checking consistency of delivery data based on spatial encoding and related device

The application discloses a delivery data consistency checking method based on space coding and related equipment, comprising the following steps: extracting attribute information related to a to-be-checked space unit from project contract data, design model data and field measurement data respectively, and uniformly associating each extracted attribute information to the same target level space identifier corresponding to the to-be-checked space unit; performing consistency checking on the to-be-checked space unit based on the attribute information associated with the target level space identifier; in response to deviation in the consistency checking, generating a checking result containing the target level space identifier and a description of the deviation, and triggering a preset disposal process.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

A computer-implemented method for generating a multimodal machine learning model, and related computer-implemented methods

A method for enhancing spatial understanding of machine learning models for at least increasing the generation capabilities of audio and image (video) models is described. The method includes collection of audio spatial data, preparing and processing of the audio spatial data, creating datasets with training inputs and outputs by some combination of the audio spatial data with other modalities, adding either encoder and-or decoder for audio spatial data to the model, training the machine learnings model using the dataset and optionally removing the audio spatial encoder and decoder. Some alternatives include the possibility to combine images with audio spatial frequency images or combine video with frequency audio-spatial temporal data.
Owner:MONAVA AB

Infrared anomaly region detection and segmentation method based on cascade architecture

This invention discloses an infrared anomaly region detection and segmentation method based on a cascaded architecture, belonging to the fields of infrared image processing and computer vision technology. Addressing the problems of high false detection rate and low segmentation accuracy in existing infrared anomaly region detection methods, this method employs a three-stage cascaded detection-filtering-segmentation architecture: After preprocessing the infrared image, it is input into a target detection model to generate candidate bounding boxes; multi-bounding box fusion technology is used to calculate spatial cue confidence; a secondary confidence assessment and false positive filtering are performed based on the channel attention mechanism and inverse residual structure of a lightweight feature extraction network, generating a filtering score; high-confidence detection boxes are converted into effective spatial codes and used as spatial cue words input into the segmentation model to complete pixel-level fine segmentation. This invention improves positioning accuracy through a multi-bounding box fusion confidence mechanism, reduces the false detection rate through secondary filtering, and achieves high-precision pixel-level segmentation without manual annotation through effective spatial code conversion.
Owner:LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY

Power prediction method and device, storage medium and electronic device

The invention discloses a power prediction method and device, a storage medium and an electronic device, and the method comprises the steps: determining a plurality of first parameter sequences according to a first multi-source parameter of a target base station in a first time period, carrying out the time coding and space coding of the plurality of first parameter sequences, so as to generate a plurality of fusion vectors, the first multi-source parameters at least comprise one of a meteorological parameter, an environmental parameter, a position parameter, a service parameter and a photovoltaic parameter, and the first time period is a time period before the current time; converting each fusion vector into an attention matrix, calculating a space-time attention output value corresponding to each fusion vector according to the attention matrix and a space-time mask mechanism, and extracting a first feature vector corresponding to each fusion vector according to each space-time attention output value, and predicting the photovoltaic power generation power and the load power of the target base station in a second time period according to the plurality of first feature vectors, wherein the second time period is a time period after the current time.
Owner:CHINA TELECOM CORP LTD