Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

159 results about "Semantic space" patented technology

Semantic spaces in the natural language domain aim to create representations of natural language that are capable of capturing meaning. The original motivation for semantic spaces stems from two core challenges of natural language: Vocabulary mismatch (the fact that the same meaning can be expressed in many ways) and ambiguity of natural language (the fact that the same term can have several ...

A large language model multilingual enhancement method and system based on model combination

This application discloses a method and system for multilingual enhancement based on a large language model using model ensemble. The system includes: a pre-trained multilingual translation model, a semantic representation mapping module, and a large language model. The multilingual translation model is used for multilingual semantic modeling and language generation, including a multilingual encoder module and a multilingual decoder module. The semantic representation mapping module is used to transform the latent space representations of different models into an interactive unified semantic space based on a cross-model representation mapping mechanism. The output of the multilingual encoder is mapped to the unified semantic representation space of the large language model, and the mapped semantics are input into the large language model to perform language-independent instruction understanding. The intermediate semantic representation output by the large language model is mapped and transformed to a cross-attention representation space, generating the final output text under the target language distribution. The system of this application outperforms existing technologies in terms of efficiency, stability, and generation quality in multilingual capability extension.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A method and system for coronary angiography image segmentation based on semantic guidance

The application discloses a coronary angiography image segmentation method and system based on semantic guidance, and belongs to the field of medical image processing. In view of the problems of semantic inconsistency, easy breaking of blood vessels and easy loss of small branches in multi-scale feature fusion of existing methods, the method of the application extracts multi-scale local structure features through a main encoder, and simultaneously extracts multi-scale original global semantic features by using an image encoder of a pre-trained medical visual basic model; after aligning the multi-scale original global semantic features, the multi-scale original global semantic features are spliced, semantic bases are constructed by performing gate fusion in a unified semantic space; multi-scale semantic reconstruction features matched with each layer of a decoder are generated and decoded by the semantic bases, and the multi-scale semantic reconstruction features are injected into the decoder layer by layer, and the multi-scale local structure features of the encoder are connected by jump connection, so that the features are guided to recover. The application realizes cross-scale semantic consistency modeling, effectively improves the blood vessel connectivity and small branch reservation capability, and is suitable for coronary angiography image segmentation.
Owner:JIANGXI AGRICULTURAL UNIVERSITY

A bearing fault diagnosis method and system based on knowledge enhancement and multi-path distillation

The present application discloses a bearing fault diagnosis method and system based on knowledge enhancement and multi-path distillation, which relates to the fields of intelligent operation and maintenance and industrial equipment health management. The method includes obtaining sensor signals and visual image data of the bearing operation; based on an asynchronous dual-channel architecture, correspondingly extracting signal features of the sensor signals and image features of the visual image data, and performing time synchronization on the signal features and the image features; using a multi-modal bottleneck Transformer module to fuse the synchronized signal features and the synchronized image features; based on a maintenance knowledge graph dynamically constructed from a bearing maintenance manual, combining a text generation model to map the fused features to a semantic space and generate a fault diagnosis report. The present application can improve the recognition accuracy, real-time performance and interpretability of diagnosis results of bearing faults.
Owner:HEFEI UNIV OF TECH

Art design feature conflict detection method and system based on deep learning

This application relates to a method and system for detecting art and design feature conflicts based on deep learning. The method includes: acquiring image data of the target artwork; extracting multi-dimensional feature vectors using a multimodal feature extraction network; generating a multi-level feature matrix through multi-scale decomposition; mapping the matrix to a preset semantic space to construct a semantic feature matrix; constructing positive and negative sample pairs using a triplet sampling strategy; calculating the cross-modal feature similarity matrix between the semantic feature matrix of the target artwork and a benchmark feature matrix library using a contrastive loss function; and generating a visualized detection report based on a similarity threshold. This technology solves the technical problems of low efficiency, strong subjectivity, and difficulty in quantifying and analyzing complex conflict relationships between design elements in traditional manual detection, achieving automated, multi-dimensional, accurate identification, and visualized presentation of art and design feature conflicts.
Owner:GUANGXI MODERN VOCATIONAL & TECH COLLEGE

A smart contract vulnerability detection model and device

PendingCN122286778AEnhancing Semantic ConsistencyEnhanced Representational CapabilitiesFeature extractionAlgorithm
This invention discloses a smart contract vulnerability detection model and device. The model includes a code structure feature extraction module, a thought chain text feature extraction module, a semantic space adaptation module, a multi-head cross-attention module, and a vulnerability classification module. Based on code structure information modeling, this invention further introduces vulnerability reasoning thought chain text features and achieves effective alignment between code structure semantics and vulnerability reasoning semantics through the synergistic effect of semantic space adaptation and multi-head cross-attention. Compared to detection methods that rely solely on overall semantic modeling of a single code modality, this invention no longer depends solely on the overall statistical features of the source code for vulnerability identification. Instead, guided by vulnerability reasoning semantics, it further focuses on core code segments strongly related to the vulnerability triggering logic, reducing the interference of irrelevant business code and redundant code on the feature extraction process, and improving the accuracy and robustness of vulnerability detection.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

Page element semantic analysis method based on visual language enhancement

The present application relates to the technical field of image understanding, in particular to a page element semantic analysis method based on visual language enhancement; a dense optical flow field and inter-frame state transition rule joint analysis is performed on an acquired visual image frame sequence, a two-dimensional spatial region in which a local rendering state is abnormal is identified, and a space-time feature matrix is constructed; a space-time feature matrix and an edge pixel brightness attenuation gradient and a visual darkening feature matrix of a neighborhood environment are extracted, and a pseudo-depth blocking parameter is identified; foreground multi-modal features inside the space-time feature matrix and background multi-modal features of an external environment are extracted, and a relative entropy divergence in a multi-modal semantic space is calculated; a multi-dimensional joint determination model is constructed using the pseudo-depth blocking parameter and the relative entropy divergence, and whether the spatial region is an interference pop-up window is identified; when the identification result is an interference pop-up window, warning data containing attribute information of the interference pop-up window is output, and an adaptive locking mechanism of interface analysis state is triggered.
Owner:ZHEJIANG FULIN TECH CO LTD

A multi-modal collaborative denoising commodity recommendation method based on modal balance

PendingCN122367584AEngineeringBack propagation algorithm
This invention discloses a multimodal collaborative denoising product recommendation method based on modality balance. The method first constructs behavior-aligned multimodal semantic encoding and projects it onto the behavior semantic space to obtain a multimodal feature product sequence. Next, for the multimodal feature product sequence, a multimodal-aware collaborative denoising module is constructed to collaboratively filter noise from each modality, resulting in a denoised intermediate product sequence. By introducing positional encoding in conjunction with the collaborative denoising module, an enhanced product sequence is obtained. Finally, based on the enhanced product sequence, cross-modal fusion weights are generated, prediction scores are calculated, and the product with the highest score is recommended. A joint loss function is constructed, and the global parameters are iteratively updated using a backpropagation algorithm. This invention effectively solves the problems of noise interference and modality learning imbalance in multimodal recommendation, suppresses the excessive dominance of strong modalities in the early stages of training, and improves the accuracy and robustness of product recommendations.
Owner:HANGZHOU DIANZI UNIV

A Method and System for Constructing Medical Knowledge Graphs Based on Artificial Intelligence and Cross-Dimensional Alignment

PendingCN122311392AMedical treatmentDatabase retrieval
This invention relates to a method and system for constructing a medical knowledge graph based on cross-dimensional alignment using artificial intelligence, belonging to the field of medical knowledge graphs. This method integrates authoritative databases from multiple medical domains by constructing a unified medical semantic space across scales and modalities; and combines deep representation learning and manifold alignment techniques to design a joint embedding and dynamic alignment framework for multi-source heterogeneous knowledge, achieving semantic unification and structure-fidelity mapping of different standard systems in a low-dimensional dense space. This invention not only effectively alleviates the semantic fragmentation problem between ontology but also significantly improves the system's semantic understanding and cross-database retrieval capabilities for complex clinical texts, providing a computable and reasonable knowledge foundation for clinical decision support, disease prediction, and precision medicine.
Owner:THE SECOND AFFILIATED HOSPITAL ARMY MEDICAL UNIV

A low-resource neural machine translation method fusing shared semantic space and bidirectional iterative generation optimization

This invention belongs to the field of natural language processing and neural machine translation technology, and discloses a low-resource neural machine translation method that integrates a shared semantic space and bidirectional iterative generative optimization (BIGO). The method includes constructing a shared semantic space module and a bidirectional iterative generative optimization module. First, this invention utilizes singular value decomposition and entropy regularization for optimal transport to map the source language and the target low-resource language into a unified semantic space, achieving high consistency between the two languages ​​at the representation level, thereby significantly improving the model's cross-language generalization ability under conditions of insufficient data. Subsequently, this invention continuously improves translation quality through alternating training of forward and backward Transformer models, in a loop of generating pseudo-bilingual corpora, backward reconstruction, and joint forward and backward optimization. This allows the model to correct errors generated in the previous iteration and enhance semantic consistency in each iteration. This invention can stably achieve dual convergence of model parameters and semantic alignment matrix, significantly improving the accuracy, robustness, and long sentence consistency of low-resource language translation, and effectively reducing training instability and overfitting.
Owner:XI'AN POLYTECHNIC UNIVERSITY

A deep learning-based multi-modal perception fusion unmanned aerial vehicle cluster control method

The present application relates to the technical field of unmanned aerial vehicle cluster control, in particular to a multi-modal perception fusion unmanned aerial vehicle cluster control method and device based on deep learning, equipment and storage medium. The present application collects gesture, voice, visual and touch data in parallel through a multi-modal perception access layer, and pre-processes and synchronizes the data. A special deep network is used to extract high-level features rich in cluster semantics. Through cross-modal semantic alignment, graph neural network correlation learning, and multi-dimensional conflict detection and intelligent resolution mechanism, the present application realizes the deep fusion of the four modes in the cluster control semantic space and the accurate analysis of the intention. The present application uses hierarchical task decomposition, virtual structure formation control and multi-objective optimization algorithm to efficiently convert high-level instructions into single-machine executable tasks and coordination strategies. The present application also performs real-time state feedback and adaptive parameter optimization, improving the accuracy, reliability, environmental adaptability and human-computer interaction efficiency of large-scale unmanned aerial vehicle cluster control.
Owner:BEIJING INST OF TECH

Cross-modal attention fusion method and device based on voiceprint features and language semantics

The application relates to the technical field of speech recognition, and discloses a cross-modal attention fusion method and device based on voiceprint features and language semantics, which comprises the following steps: performing a feature extraction operation on speech data to obtain voiceprint features; based on a sentiment perception mask mechanism and a context sentiment memory unit, fusing sentiment features into text data to obtain semantic features; constructing a cross-modal relationship graph of the voiceprint features and the semantic features; determining a time sequence dependency relationship between the voiceprint features and the semantic features based on the edge weights generated by the cross-modal relationship graph; projecting the voiceprint features and the semantic features into the same semantic space based on the time sequence dependency relationship, and performing a feature reconstruction fusion operation based on a reconstruction loss function to obtain voiceprint and semantic fusion features; and adjusting a dialogue strategy determined based on the voiceprint and semantic fusion features; and using the adjusted dialogue strategy to generate interactive response data with historical interactive data, so that a high-accuracy and personalized speech interaction experience is realized.
Owner:GUANGDONG GUANGXIN COMM SERVICES COMPANY

Pet status determination methods, interaction methods, devices, electronic devices and media

This invention discloses a pet state determination method, interaction method, device, electronic device, and medium. The pet state determination method includes: acquiring multimodal data of the pet, wherein the multimodal data includes at least two of video data, audio data, behavioral posture data, and physiological data; performing feature extraction and multimodal fusion on the multimodal data to generate a pet feature vector representing the pet's state; mapping the pet feature vector to a pre-constructed pet state semantic space to obtain a pet state embedding vector; and generating human-understandable state description text matching the pet's current state based on the pet state embedding vector and a preset natural language generation model. This invention can improve the accuracy of pet state determination.
Owner:SHENZHEN KOLAMAMA TECH CO LTD

A physical perception time-frequency semantic alignment timing prediction method and system for edge embedded devices

The application discloses a physical perception time-frequency semantic alignment timing prediction method and system for edge embedded devices, and belongs to the technical field of deep learning and industrial artificial intelligence. The method comprises the following steps: preprocessing historical timing data to obtain a standardized input sequence; then decoupling the sequence into a low-frequency trend component and a high-frequency detail component through causal stationary wavelet transform, so that feature extraction only depends on historical observation data; then inputting the decoupled components into a context perception prototype reprogramming module, projecting continuous physical features into a discrete semantic space of a frozen large language model by using a mixed semantic prototype library and a double-flow cross-attention mechanism; then dynamically fusing the trend and detail components through a semantic gating network to generate fused semantic features; and finally obtaining a final prediction result after splicing the fused semantic features and wavelet enhancement prompts. The application realizes high-precision and low-delay timing prediction.
Owner:INST OF ELECTRICAL ENG CHINESE ACAD OF SCI

A Deep Learning-Based Intelligent Service Robot Control System

This invention relates to the field of intelligent robot control technology and provides a deep learning-based intelligent service robot control system. The system constructs a multimodal semantic reasoning model to fuse and model multi-source perception information such as vision, LiDAR, and radar. Within a unified semantic space, it performs tasks such as target detection, navigable area understanding, and dynamic obstacle analysis, improving the consistency of environmental perception. To address the training instability caused by non-stationary changes in the indoor environment, a smooth function-based two-layer optimization method is introduced to achieve adaptive optimization of multi-task trade-offs and fusion strategies, enhancing the stability and robustness of semantic reasoning. Furthermore, at the decision-making and control level, a hierarchical reinforcement learning method based on an enhanced Option-Critic architecture is employed to achieve synergy between task decision-making and continuous control, thereby improving the robot's safety and execution efficiency in complex service tasks.
Owner:TAISHAN UNIV

Integrated Image Restoration Method and System

This invention provides an integrated image restoration method and system, belonging to the field of image processing. The method includes constructing a text-guided integrated image restoration model (TGMIR), which maps text prompts to a semantic space consistent with image features and achieves hierarchical semantic control at three levels: channel attention, spatial attention, and cross-modal collaborative attention. The image is then restored using the TGMIR. This invention aims to overcome the core limitations of existing integrated image restoration models under multi-degradation conditions, including insufficient degradation semantics, significant cross-degradation interference, weak modal collaboration capabilities, and insufficient feature fusion.
Owner:SOUTHWEST PETROLEUM UNIV

Power load scheduling method and device based on large language model

The application discloses a power load scheduling method and device based on a large language model, relates to the technical field of power scheduling, and mainly aims to solve the problem of poor accuracy of existing power load scheduling. The method comprises the following steps: acquiring scheduling text data of power load, wherein the scheduling text data comprises load time sequence data, meteorological text data, equipment log data and power grid index data; performing feature extraction on the scheduling text data to obtain multi-modal features, and encoding the multi-modal features into a semantic space to obtain multi-modal feature encoding; and performing scheduling prediction processing on the multi-modal feature encoding after consistency processing based on a large language model after model training to obtain a power load scheduling result.
Owner:STATE GRID TIANJIN ELECTRIC POWER COMPANY +1

Image semantic information generation method and electronic device

The application discloses an image semantic information generation method and an electronic device, and relates to the technical field of artificial intelligence. The method comprises the following steps: generating multi-dimensional image semantic features according to a target subject and context information of an image to be processed; inputting the multi-dimensional image semantic features into a semantic label graph and a nonlinear mapping network respectively to obtain semantic representation information and a semantic space alignment prediction result with the same dimension as a label number; and fusing the semantic space alignment prediction result and the semantic representation information to generate image semantic information of the image to be processed. The semantic label graph is a heterogeneous graph neural network with semantic labels as graph nodes, label semantic description features as graph node features, and node connection edges determined based on semantic similarity and co-occurrence relationships between the semantic labels. The application can solve the problem that related technologies cannot meet the demand of users for the quality of semantic information generation, and can generate high-precision image semantic information.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Intelligent digital textbook generation method and system based on generative artificial intelligence

The application provides a wisdom digital teaching material generation method and system based on generative artificial intelligence, relates to the technical field of generative artificial intelligence, and first collects an initial subject knowledge corpus set and a preset cognitive gradient evolution template, performs latent semantic space mapping and knowledge atomization cutting processing on the corpus set, is reconstructed into a basic cognitive atomic unit, and is given an encoding mark and extracts a semantic feature vector. Then, a cognitive conflict network between the basic cognitive atomic units is constructed, a pre-trained model is called to perform text content deepening and cognitive guidance question generation processing driven by cross-concept thinking conflict, and a strengthened cognitive atomic unit set is obtained. Finally, dynamic cognitive path planning is performed according to the strengthened cognitive atomic unit set and the cognitive conflict network, a wisdom digital teaching material file with a nonlinear thinking transition path is generated, the personalized learning needs of students are met, and the learning effect is improved.
Owner:BEIJING VOCATIONAL COLLEGE OF ECONOMICS & MANAGEMENT (BEIJING MANAGER COLLEGE) +1

A Method and System for Intelligence Generation Based on Feature Fingerprint Storage and Spatiotemporal Geometric Correction

PendingCN122313306ASmooth deploymentReduce video memory usageData streamEngineering
This invention discloses an intelligence generation method and system based on feature fingerprint storage and spatiotemporal geometric correction, belonging to the field of intelligent remote sensing image processing technology. The method involves: accessing heterogeneous data streams generated from multi-source remote sensing images in a synchronized time sequence; inputting a shared backbone network and a modality adaptation layer to generate a current feature stream in a unified semantic space; asynchronously retrieving historical baseline features from an in-memory feature fingerprint database based on geographic coordinates; fusing satellite imaging parameters with the current feature stream and inputting it into a spatial transformation network, resampling historical baseline features using a regression transformation matrix to achieve spatiotemporal geometric correction in the feature domain; weighted fusing of the current feature streams from each modality to generate a fused feature stream; and parallel interpretation and semantic encapsulation of the fused feature stream to generate a natural language intelligence report. This invention reduces memory usage and data throughput, shortens the intelligence generation cycle, reduces the false alarm rate, and achieves synergistic optimization of timeliness, accuracy, and automation.
Owner:XIAN HUIGUANG RIXIN OPTOELECTRONICS TECHNOLOGY CO LTD

A gpu-based multi-modal indexing integration method

The application discloses a GPU-based multi-modal index integration method. The method comprises the following steps: acquiring multi-modal original data and extracting a feature vector; mapping the feature vector to a unified common semantic space; constructing a double-layer hybrid index structure on a GPU, wherein a first layer adopts IVF-PQ index for coarse screening, and a second layer adopts graph index for fine screening; after receiving a query request, performing two-stage parallel search on the GPU, first screening a candidate set through IVF-PQ index, and then fine screening a final result through graph index. The application solves the cross-modal semantic alignment problem through the unified semantic space, realizes high-precision and low-delay retrieval of massive multi-modal data by using GPU parallel computing and a double-layer index structure, and effectively overcomes the technical bottleneck that precision and efficiency are difficult to be considered in the prior art.
Owner:FENGHE SMART TECH (SHANGHAI) CO LTD

A bilingual question-answering method based on collaborative training of knowledge retrieval and generation

The application discloses a bilingual question and answer method based on knowledge retrieval and generation collaborative training, and belongs to the technical field of intelligent question and answer, and comprises the following steps: obtaining a corpus g0 and a question and answer library; mapping sentences in the g0 to the same semantic space to form a set A; performing regular hierarchical clustering on the set A to generate a semantic prototype vector of each clustering cluster; generating a knowledge graph g1 and a causal diagram g2 to form a retrieval source set G; marking a sample set and a main retrieval source for a question q of the question and answer library; constructing a knowledge retrieval enhancement model and training the knowledge retrieval enhancement model into a bilingual question and answer model, which is used in intelligent question and answer. The application introduces hierarchical semantic aggregation, cross-source adaptive gating and causal biasing retrieval mechanism, solves existing problems such as scattered knowledge, language alignment and self-learning update, provides a new technical path for low-resource cross-language question and answer, and greatly improves the quality and application value of a cross-language knowledge enhancement question and answer system.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY +1

Data processing method and device, electronic equipment, storage medium and program product

PendingCN122313358AVideo processingEngineering
This disclosure provides a data processing method, apparatus, electronic device, storage medium, and program product, relating to the field of data processing technology, and particularly to large-scale modeling, artificial intelligence, and video processing technology. The implementation scheme includes: acquiring image and text information corresponding to the video to be processed; encoding the image and text information respectively to obtain image features and text features of the video to be processed; fusing the image and text features to obtain a multimodal representation of the video to be processed; mapping the multimodal representation to obtain risk identification features, wherein the mapping is used to map the multimodal representation from a general semantic space to a risk identification semantic space, where the dimension of the risk identification semantic space is lower than that of the general semantic space; and processing the risk identification features using at least two classification heads to obtain the risk identification result of the video to be processed.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Cross-modal attention fusion method and device based on voiceprint features and language semantics

The application relates to the technical field of speech recognition, and discloses a cross-modal attention fusion method and device based on voiceprint features and language semantics, which comprises the following steps: performing a feature extraction operation on speech data to obtain voiceprint features; based on a sentiment perception mask mechanism and a context sentiment memory unit, fusing sentiment features into text data to obtain semantic features; constructing a cross-modal relationship graph of the voiceprint features and the semantic features; determining a time sequence dependency relationship between the voiceprint features and the semantic features based on the edge weights generated by the cross-modal relationship graph; projecting the voiceprint features and the semantic features into the same semantic space based on the time sequence dependency relationship, and performing a feature reconstruction fusion operation based on a reconstruction loss function to obtain voiceprint and semantic fusion features; and adjusting a dialogue strategy determined based on the voiceprint and semantic fusion features; and using the adjusted dialogue strategy to generate interactive response data with historical interactive data, so that a high-accuracy and personalized speech interaction experience is realized.
Owner:GUANGDONG GUANGXIN COMM SERVICES COMPANY

A Deep Multi-View Clustering Method Based on Multi-Granularity Contrast Learning

This invention provides a deep multi-view clustering method based on multi-granularity contrastive learning, belonging to the field of computer vision technology. It solves the technical problem that single-granularity feature representations are insufficient to comprehensively capture the multi-level semantic information of data and the quality differences between different views. The method includes the following steps: S10, designing an independent encoder for each view to extract latent embedding features; S20, mapping the attention-enhanced features of each view to the same semantic space; S30, dynamically assigning weights based on the similarity between the features of each view and the global representation; S40, granular-level contrastive learning can enhance the compactness of the cluster structure within the granular sphere and align the granular sphere structure between cross-view granular spheres; S50, performing K-means clustering on the final global features to obtain the final prediction result. This invention enhances the weight of key features by adding a feature attention enhancement module to the view features.
Owner:NANTONG UNIV

A patent retrieval and technology monitoring method for multi-source heterogeneous data

PendingCN122285845AHigh recognition sensitivityImprove noise immunitySemantic representationEngineering
This invention relates to a method for patent retrieval and technology monitoring of multi-source heterogeneous data. It constructs a multimodal semantic representation including text, images, and time-series fragments, and generates a joint temporal representation vector by combining attention perturbation entropy and term evolution distance. A unified semantic space is established based on historical technology anchors, dynamically monitoring and quantifying semantic shifts to achieve real-time determination of drift trends. When substantial semantic drift occurs, local semantic correction is performed using gated contrastive learning and an anchor-aware adaptation layer to guide the orderly alignment of newly added modal features, ensuring the continuation and expansion of the original semantic topology. Simultaneously, an anchor decay and candidate new anchor generation mechanism is designed, combining cross-validation from patent citations, industry reports, and policy documents to achieve incremental expansion of the semantically stable anchor set. This scheme effectively improves the fusion accuracy, real-time adaptability, and dynamic topology update capability of multimodal patent data.
Owner:SHANDONG BITGO DATA TECHNOLOGY CO LTD

Systems and methods for mapping a term to a vector representation in a semantic space

ActiveUS12675493B2AlgorithmSemantic space
A method and system is provided for mapping a term to a vector representation in a semantic space. Provided techniques allow for efficient and accurate determination of vector representations for query terms that are terms of emerging interest or are otherwise not included in a set of terms for which vector representations are pre-calculated.
Owner:NFERENCE INC

A novel multi-modal semantic-spatial representation method for embodied perception and spatial reasoning

PendingCN122334280AEmbodied perceptionSpatial perception
This invention discloses a novel multimodal semantic-spatial representation method for embodied perception and spatial reasoning, belonging to the field of embodied perception. Addressing the two major pain points of existing methods—insufficient spatial coherence in semantic reasoning and poor relation consistency and interpretability—this invention designs an optimization mechanism: on the one hand, it designs a transitive relation learning mechanism to correct relation conflicts and missing relations in reasoning under weak supervision, improving the completeness and logical consistency of spatial representation; on the other hand, relying solely on RGB images and depth maps, it integrates monocular depth estimation as a spatial modality into embodied perception, injecting implicit three-dimensional spatial attributes into the semantic structure, constructing a unified multimodal semantic-spatial representation, and enhancing the spatial coherence of semantic reasoning. This invention improves the performance of structured scene modeling and spatial grounding description generation through advanced spatial representation, and can be widely used in embodied spatial perception scenarios.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Method and apparatus for multi-modal data alignment and augmentation in industrial large model training

PendingCN122332948AScale modelPhysical reality
This specification provides a method and apparatus for multimodal data alignment and enhancement in industrial large-scale model training. By integrating an industrial knowledge graph with physical law constraints into a cross-modal semantic distillation architecture, heterogeneous features such as vibration, vision, and process parameters are aligned in a semantic space that conforms to engineering laws, effectively solving the problem of fragmented semantic expression and enhancing the depth of understanding of industrial processes. By introducing a dual-discriminator adversarial enhancement strategy that includes a realism discriminator and a physical consistency discriminator, and combining the constraints of an industrial anomaly pattern library, it ensures that the generated data satisfies both statistical distribution characteristics and physical reality, significantly reducing the data distortion rate. This provides high-fidelity training data for industrial large-scale models, ultimately achieving a comprehensive improvement in model convergence speed, anomaly detection accuracy, and decision credibility, meeting the application requirements of high precision and high real-time performance in industrial scenarios.
Owner:BEIJING EASY TIMES DIGITAL TECH

A large model generated content tamper-resistant watermarking method based on semantic steganography

The application discloses a large model generated content tamper-resistant watermarking method based on semantic steganography, which comprises the following steps: extracting core semantics from an original text to generate a semantic fingerprint, and performing error correction coding on the semantic fingerprint to obtain a watermark bit sequence; dividing the text into multiple text blocks; through iteration, generating a semantic space geometric partition for each text block according to a secret key and the semantic history of all previously processed blocks through a pseudo-random process; and using a large language model to perform constraint rewriting on the text block, so that the semantic vector of the rewritten text block falls into the geometric partition specified by the current watermark bit, until all bits are embedded. The application realizes strong robustness against interpretation attacks by implanting the watermark into the semantic space, and based on the context-dependent mechanism of the hash chain, any tampering will cause a chain of errors, thereby locating the tampered position, and providing an effective technical means for integrity verification and copyright protection of the large model generated content.
Owner:NANJING YIZHENG COMM INFORMATION TECH CO LTD