Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2948 results about "Semantic feature" patented technology

Semantic features represent the basic conceptual components of meaning for any lexical item. An individual semantic feature constitutes one component of a word's intension, which is the inherent sense or concept evoked. Linguistic meaning of a word is proposed to arise from contrasts and significant differences with other words. Semantic features enable linguistics to explain how words that share certain features may be members of the same semantic domain. Correspondingly, the contrast in meanings of words is explained by diverging semantic features. For example, father and son share the common components of "human", "kinship", "male" and are thus part of a semantic domain of male family relations. They differ in terms of "generation" and "adulthood", which is what gives each its individual meaning.

Large language model construction method fused with spatial semantic understanding

The invention relates to a large language model construction method and system fused with spatial semantic understanding, and the method comprises the steps: obtaining a multi-source heterogeneous corpus, and extracting an entity, an attribute and a business rule; extracting a semantic feature vector set based on the multi-source heterogeneous corpus, and constructing an entity relationship network and an enhanced knowledge graph; generating an enhanced training sample, and training the general large language model to obtain a primary large language model; generating a verification sample set and performing verification; identifying a specific weakness pattern, and generating a corresponding confrontation sample and a knowledge enhancement sample; training the primary large language model to obtain an optimized large language model; in conclusion, the enhanced knowledge graph fusing the spatial semantic features and the business rules is constructed, and the gradient training samples are generated based on the graph to perform multi-stage model training and optimization, so that the method has the effects of improving the internalized understanding ability of the model for the spatial semantics and the business rules and enhancing the reliability of multi-step spatial reasoning.
Owner:URBAN PLANNING & DESIGN INST OF SHENZHEN UPDIS

Visual language navigation method for cross-modal alignment in dynamic shielding environment

The invention discloses a visual language navigation method for cross-modal alignment in a dynamic shielding environment, and the method comprises the steps: collecting multi-modal data through a visual sensor, an inertial measurement unit, a laser radar and the like, and carrying out the preprocessing and time synchronization; sensing the dynamic shielding object through a model composed of a convolutional neural network and a long-short-term memory network, and estimating the future change of the dynamic shielding object in combination with a space-time sequence prediction algorithm; a double-branch convolutional neural network and a Transform based on a dynamic attention mechanism are adopted to respectively extract visual and semantic features and fuse the visual and semantic features; on the basis of occlusion prediction, potential occlusion region features are extracted in advance from a time dimension, an occluded image is repaired by using a generative adversarial network and geometric constraints in a space dimension, and cross-modal feature alignment is optimized through an attention mechanism; planning a path by using a hybrid reinforcement learning algorithm based on a deep Q network-space and a fast exploration random tree, and dynamically adjusting according to real-time shielding; according to the method, the accuracy, adaptability and reliability of visual language navigation in a dynamic shielding environment are improved.
Owner:SHANGHAI JIAOTONG UNIV

Multi-modal remote sensing semantic segmentation method and system for learning frequency domain fusion

The invention discloses a multi-modal remote sensing semantic segmentation method and system for learning frequency domain fusion. The method comprises the following steps: respectively extracting multi-scale features of two modal input images by adopting a double-branch encoder; sequentially executing frequency domain decoupling and fusion, mutual information constraint-based feature optimization and low-frequency guided cross-modal fusion processing on each scale feature to generate a fused semantic feature; and performing up-sampling and feature refining on the fused features through a decoder, and outputting a full-resolution segmentation prediction map. According to the multi-modal remote sensing image semantic segmentation method, modal sharing information and specific details are effectively separated through frequency domain decoupling, feature representation is optimized through mutual information constraint, adaptive feature fusion is achieved in combination with an attention mechanism, and the accuracy and robustness of multi-modal remote sensing image semantic segmentation are remarkably improved.
Owner:NORTHEAST FORESTRY UNIV

Illegal content auditing method and device based on multi-modal data, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical treatment and health and the like, and discloses a violation content auditing method, device and equipment based on multi-modal data and a medium. Inputting the visual semantic features and the composite audio features into a multi-modal model, generating fusion features through model alignment and fusion, and analyzing the fusion features based on a knowledge base to judge whether illegal content fragments exist in the multi-modal data, and when the illegal content fragments exist, positioning the illegal content fragments in the multi-modal data and generating an auditing report. According to the method, the visual semantic features and the composite audio features are fused, cross-modal compliance analysis is realized in combination with the knowledge base, frame-level or time-axis-level positioning is performed on the illegal content segments, and the auditing report containing the evidence is generated, so that the problems of insufficient single-modal detection accuracy and poor positioning capability are solved, and the auditing accuracy is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Abnormal scene detection method based on visual and semantic feature fusion

The invention discloses an abnormal scene detection method based on visual and semantic feature fusion, and relates to the technical field of safety monitoring and intelligent identification, and the method comprises the steps: carrying out the preprocessing of a collected original image, and obtaining a preprocessed image; forming a multi-modal input pair by the preprocessed image and a predefined structured prompt statement; inputting the multi-modal input pair into the visual language large model, and outputting semantic features including visual feature vectors and text vectors; inputting the preprocessed image into a target detection model, and outputting visual features; fusing the semantic features and the visual features through a cross-modal attention mechanism to obtain multi-scale fusion features; and inputting the multi-scale fusion features into detection heads of all scales, executing abnormal scene detection, and outputting an abnormal detection result. When the unconventional object is identified in the abnormal scene, the visual features and the semantic features are fused to perform abnormal scene detection, so that the strong perception capability of a complex scene is realized, and false alarm or missing alarm is effectively avoided.
Owner:CHONGQING UNIV OF ARTS & SCI

Text-driven CAD modeling method and system based on diffusion and visual language model

The invention relates to the technical field of computer aided design, in particular to a text-driven CAD modeling method and system based on a diffusion and visual language model.The method comprises the steps that natural language text description is obtained, and CAD semantic features of the natural language text description are extracted; carrying out geometric standardization on the CAD semantic features by adopting a fine-tuning diffusion model, and generating a CAD view image conforming to engineering specifications; carrying out fusion by adopting a fine-tuned visual language model to generate a parameterized CAD construction sequence; a three-mode alignment mechanism is adopted, and the semantic consistency of the CAD semantic features, the CAD view images and the CAD construction sequences is checked; performing verification and post-processing on the CAD construction sequence, and outputting an executable Python code or STEP file; the CAD modeling method disclosed by the invention performs explicit modeling based on flexible modal description, and has the characteristics of high geometric constraint and high usability.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Small target detection and state perception method based on multi-scale feature fusion

The invention discloses a small target detection and state perception method based on multi-scale feature fusion, and belongs to the field of computer vision and deep learning. Multi-scale semantic features are extracted through a backbone network; two uplink fusion paths and two cascaded downlink enhancement paths are constructed, and multi-scale feature fusion is performed, so that the perception capability of targets with different sizes is enhanced, and the accuracy and robustness of detection are improved; and meanwhile, a regional state sensing mechanism is constructed based on a detection result, continuous monitoring and intelligent analysis of target space distribution, behavior trend and dynamic change are realized, and the adaptability and response speed of the system in a complex environment are improved. The method gives consideration to the detection precision and the calculation efficiency, and is suitable for real-time application scenes with limited resources.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Natural language to low code conversion method based on multi-modal reinforcement learning

The invention discloses a method for converting a natural language into a low code based on multi-modal reinforcement learning, which comprises the following steps of: performing word segmentation, embedding and multi-layer feature extraction on a natural language instruction input by a user, combining a self-attention mechanism and a graph attention mechanism, extracting and optimizing an original semantic feature vector, and automatically identifying a business field; semantic relationship triples in the domain knowledge graph are fused, and a multi-modal semantic alignment and context enhancement strategy is adopted, so that the accuracy and stability of semantic representation are remarkably improved; when semantic drift or ambiguity is detected, a multi-candidate correction mechanism is triggered to obtain a better analysis result, and the robustness of semantic understanding is enhanced; and finally, mapping an analysis result into an instruction which can be identified by a low-code platform, automatically generating a code structure, continuously optimizing a semantic model and a knowledge graph based on user feedback, realizing adaptive learning, and improving the conversion efficiency and quality from a natural language to codes.
Owner:GUANGZHOU ZHUORUI DIGITAL TECHNOLOGY CO LTD

Intelligent agent memory indexing method and system based on intention recognition

The embodiment of the invention provides an intelligent agent memory indexing method and system based on intention recognition. The method is applied to the technical field of artificial intelligence and comprises the steps of obtaining real-time question-answer data, and performing preliminary intention classification on the real-time question-answer data by utilizing a domain knowledge rule library; extracting a structured description from the real-time question and answer data after the preliminary intention classification, performing deep intention analysis in stages, and outputting a standardized intention description text; according to the standardized intention description text, acquiring an Agent operation context, performing multi-dimensional retrieval to obtain an adaptive strategy, executing the adaptive strategy, and returning a strategy evaluation result; according to a strategy evaluation result, carrying out microscopic feedback and macroscopic feedback to update a strategy library; the Agent operation context is obtained through the following steps that semantic features of a standardized intention description text are captured, and the Agent operation context corresponding to the deep semantic features is recorded based on a fine-grained metadata labeling system. According to the invention, a complete closed loop from intention identification to strategy multiplexing to strategy optimization is realized.
Owner:TERMINUSBEIJING TECH CO LTD

Cement equipment maintenance decision-making method and device based on knowledge graph and large model reasoning

The invention provides a cement equipment maintenance decision-making method and device based on a knowledge graph and large model reasoning, relates to the field of cement industry intelligent operation and maintenance, and solves the technical problem of decision-making response delay caused by knowledge fragmentation. The method comprises the following steps: extracting real-time characteristics from vibration spectrum signals, temperature curves and torque waveform data collected by an edge gateway, and extracting a work order entity triple from a natural language work order text of an EAM system; based on an equipment BOM list, a historical maintenance record and an FMEA analysis table, physical assembly constraint conditions are defined through ontology modeling to generate a cement equipment topological relation and a fault rule chain, and a knowledge graph is created to output a fault rule base with confidence coefficient weights. And inputting the real-time feature vector and the work order entity triple into a multi-modal collaborative inference engine, triggering a matched fault rule chain by combining real-time features and semantic features, outputting a fault root cause and an associated maintenance strategy ID, and labeling a logic chain. And activating the associated maintenance strategy ID and obtaining the real-time characteristic deviation degree of the maintenance strategy ID, quantifying the decision credibility through a tracing rule matching path, obtaining an executable maintenance instruction packet with a logic chain, executing the maintenance instruction packet and dynamically updating the knowledge graph based on a maintenance result. The method is used in the maintenance decision-making process of the cement equipment.
Owner:HEFEI CEMENT RESEARCH AND DESIGN INSTITUTE CO LTD

Employment information matching method and system based on data analysis

The invention discloses an employment information matching method and system based on data analysis, and particularly relates to the field of employment matching, and the method comprises the steps: collecting structured and unstructured data, and carrying out the semantic feature extraction through employing a BERT model and BiLSTM-CRF; in the preprocessing stage, entity standardization is realized through a knowledge graph, and a job seeker portrait and post model including a skill matrix and an occupational development trajectory is constructed; in the feature engineering stage, extracting four core features of skill matching degree, salary expectation integrating degree, commuting tolerance and occupational development goodness of fit; the salary integrating degree is quantitatively evaluated through a bidirectional tolerance model, a random forest model with time decay is used for dynamic weight, feature weight is generated based on historical successful cases, and a real-time feedback mechanism is introduced to adjust a weight coefficient; finally, the matching degree function fuses the weighted features and the industry trend factors, and the model is continuously optimized through a three-level updating mechanism.
Owner:BEIJING ZHONGZIHAIWAI CONSULTATION CO LTD

Document content extraction method and system based on multimodal model collaboration, terminal and medium

The invention belongs to the technical field of document content extraction, and particularly discloses a document content extraction method and system based on multimodal model collaboration, a terminal and a medium. Comprising the following steps: identifying the type of an input to-be-processed document, and judging the document type; on the basis of the type identification result, calling a multi-modal model to analyze the document content, and outputting space coordinates, visual features and semantic features of document elements; generating a content sequence according with a reading habit through a semantic sequence reconstruction algorithm; paragraph boundary detection, paragraph recombination and semantic association modeling of charts and texts are completed based on the multilayer attention network and the graph neural network; grammar error correction, format optimization and title hierarchy generation are carried out by using a large language model and a hierarchical classification network; and converting the identification result into a structured output file. According to the method, the processing requirements of different types of documents can be considered, and high-precision analysis and efficient output are realized under the scenes of complex layouts, multiple languages and formula tables.
Owner:TUOSI (SHANDONG) INFORMATION TECHNOLOGY CO LTD

Ring main unit inspection robot autonomous navigation method and system based on SLAM

The invention discloses a ring main unit inspection robot autonomous navigation method and system based on SLAM, particularly relates to the technical field of robot autonomous navigation and intelligent inspection, and is used for solving the problem of positioning drift caused by repeated features of an existing ring main unit scene. Semantic feature analysis and topological constraints are introduced into an SLAM processing flow, acquired image data and point cloud data are processed through a deep learning model, objects such as an electrical cabinet, a corridor channel and a cable trench are identified, and a semantic feature set with category labels and spatial position information is generated; and constructing a topological graph containing node spacing, connectivity and directivity constraints based on the semantic features, adding the topological graph as a constraint factor into SLAM back-end optimization, and performing joint optimization in combination with vision, a laser odometer and inertial prior information, thereby avoiding only depending on repeated geometric feature positioning, and improving the positioning accuracy. The problems of loopback misjudgment and drifting caused by feature confusion are reduced, and the pose resolving stability in the ring main unit environment is improved.
Owner:STATE GRID HUBEI ELECTRIC POWER CO XIAOGAN POWER SUPPLY CO

Intelligent AI semantic annotation method and system based on multi-modal analysis

The invention relates to the field of artificial intelligence, and discloses an intelligent AI semantic annotation method and system based on multi-modal analysis, images, audios, texts and multi-modal data of a sensor are synchronously collected through AI, deep semantic features of all modals are extracted by using a deep learning model after preprocessing, vectors are generated, and the multi-modal data are subjected to semantic annotation; fusing into a semantic scene vector by means of a cross-modal attention mechanism, constructing a cross-modal association graph, and learning a causal logic path by using a graph neural network; constructing a spatiotemporal causal diagram by combining spatiotemporal information, and integrating an event time sequence and spatial dependence by using a time sequence diagram neural network to form a venation path; according to the method, the two knowledge maps are fused into a comprehensive knowledge map, preliminary annotations are generated through rule base reasoning, a high-precision result is output after optimization, the problems of semantic loss and annotation fragmentation caused by modal splitting and context missing in a traditional method can be effectively solved, meanwhile, the manual verification cost is reduced, and the integrity and efficiency of semantic annotation in a complex scene are improved.
Owner:GUANGDONG SHUNNENG CONSTR CO LTD

Intelligent early warning analysis method for fund flow abnormity

The invention provides a fund flow abnormity intelligent early warning analysis method, and relates to the technical field of data processing, and the method comprises the steps: obtaining transaction data of a target account in a continuous time interval; semantic feature extraction is carried out, and a fund semantic unit set is generated according to a transaction initiator, a transaction receiver and a business type; performing causal correlation calculation on adjacent fund semantic units to determine behavior causal strength; when the behavior causal intensity is lower than a preset causal threshold value, marking the corresponding time period as a behavior fault region; performing context reconstruction on the account set in the behavior fault area; executing mode deviation analysis, and identifying fund flow modes with abnormal aggregation or reverse backflow characteristics; performing self-learning updating on the abnormal threshold weight of the initial judgment model to obtain a corrected judgment model; detecting subsequent transaction data according to the corrected judgment model, and generating fund flow abnormity intelligent early warning information; according to the invention, autonomy and accuracy of fund flow abnormity intelligent early warning analysis are improved.
Owner:HANGZHOU HUAYI ZHILIAN TECHNOLOGY CO LTD

RAG-based voucher classification method, medium and equipment

The invention relates to an RAG-based voucher classification method, a medium and equipment, and the method comprises the steps: receiving digital image data of a to-be-classified voucher, carrying out the multi-modal optical character recognition processing to generate a structured OCR result, extracting a text semantic feature vector and a visual layout feature vector based on the structured OCR result, carrying out the fusion of the text semantic feature vector and the visual layout feature vector to generate a multi-modal query vector, and carrying out the classification of the to-be-classified voucher. Similar samples and semantic similarity scores and category metadata thereof are obtained through approximate nearest neighbor retrieval, after an initial candidate category list is generated, key field values are extracted for each candidate category, evidence credibility scores are calculated, comprehensive confidence scores are generated by fusing the semantic similarity scores and the evidence credibility scores, reordering is conducted, and a candidate category list is obtained. And finally, selecting a classification decision path according to the score distribution, and outputting a classification result and an interpretability report. The accuracy and robustness of voucher classification are effectively improved, and complex voucher scenes with changeable formats and fuzzy semantics can be processed; and the interpretability and reliability of the classification decision are enhanced.
Owner:FUJIAN BOSS SOFTWARE

Long text intelligent review and prediction method fusing dynamic knowledge evolution mechanism

The invention provides a long text intelligent review and prediction method fusing a dynamic knowledge evolution mechanism, and relates to the technical field of text review, and the method comprises the steps: carrying out the structural analysis of a long text, constructing an initial knowledge graph, and generating an evolution knowledge graph through combining a time sequence change mode of an entity relationship in a historical text; calculating semantic similarity between word vectors and graph embedding to realize information interaction; deep semantic features are extracted to calculate the mahalanobis distance between the deep semantic features and an abnormal category prototype to determine an abnormal mode; and combining historical evolution trajectory modeling time sequence characterization to predict an abnormal development trend. And the accuracy and prediction capability of long text review can be effectively improved.
Owner:BEIJING FEIRUI XINGTU TECH CO LTD

Image review method fusing semantic comprehension and visual identification

The invention belongs to the technical field of machine room safety monitoring, and discloses an image review method fusing semantic understanding and visual recognition, which comprises the following steps: calculating visual / semantic feature dynamic credibility in real time through an exponential weighted moving average algorithm in combination with environment interference and equipment state parameters; resNet50 is adopted to extract visual features in global and local branches, and a BERT model is adopted to encode semantic features; correcting the feature correlation degree according to the scene, calculating a dynamic weight, carrying out heterogeneous calibration through an attention mechanism, and calling priority rules such as'physical security features are higher than behavior features' for arbitration during conflicts; the method is advantaged in that low-credibility feature interference fusion is avoided, abnormity identification accuracy in a complex scene of a machine room is improved, multi-modal data cooperation demands are adapted, scene labels are marked based on time, work orders and historical data dimensions, an exclusive sub-model is constructed for a high-density scene through DBSCAN density clustering, and low-frequency scene parameters are migrated to a similar model.
Owner:QINGYUN CLOUD COMPUTING (SHENZHEN) CO LTD

File retrieval method, device and equipment based on multi-modal AI and storage medium

The invention discloses a file retrieval method, device and equipment based on multi-modal AI and a storage medium, relates to the technical field of data retrieval, and aims to solve the problem that traditional file retrieval is low in efficiency and precision. The method comprises the following steps: performing content analysis on a multi-modal file including a picture file, a video file and a document file of text and / or visual information, and extracting a text semantic feature, a picture visual feature and a video time sequence feature; mapping the text semantic feature, the picture visual feature and the video time sequence feature to a cross-modal semantic space used for representing a high-dimensional vector associated with different modal features, and generating a cross-modal association vector; according to the modal type of the retrieval information, a corresponding retrieval module in a mixed retrieval engine is called to conduct retrieval in a cross-modal semantic space, an initial retrieval result is obtained, and the mixed retrieval engine comprises a text retrieval module, a visual retrieval module and a cross-modal fusion module; and performing dynamic weight distribution sorting on the initial retrieval result, and outputting a target matching result.
Owner:SHANXI XINDINGCHEN TECH CO LTD

Cross-modal document information extraction method based on space-semantic alignment

The invention relates to a cross-modal document information extraction method based on space-semantic alignment, and belongs to the field of artificial intelligence, computer vision and natural language processing. According to the method, the spatial feature and semantic information bidirectional alignment model is designed, by constructing the spatial feature and semantic feature bidirectional alignment model, the document layout information can dynamically adjust attention distribution of text semantic features, meanwhile, semantic information reversely optimizes the spatial features, collaborative modeling of spatial layout and semantic information is achieved, and the document layout efficiency is improved. Therefore, the accuracy and robustness of complex document information extraction are improved. According to the method, a hierarchical cross-modal information extraction model is designed, through the hierarchical cross-modal information extraction model, the overall structure of a document is recognized on the global level, local key content is focused on the regional level, fine modeling is conducted on fine-grained texts and visual elements on the entity level, and accurate recognition of a cross-modal entity and the semantic relation of the cross-modal entity is achieved; and the generalization ability and applicability of information extraction are enhanced.
Owner:BEIJING INST OF COMP TECH & APPL

Camouflage target detection method based on feature selection attention and frequency domain edge guidance

The invention discloses a camouflage target detection method based on feature selection attention and frequency domain edge guidance. According to the method, four-level features of a camouflage target image are extracted through a backbone network SMT and are respectively screened; the high-level features are input into a semantic information supplement module, and after semantic features are enhanced, the high-level features and the trunk features are sent into a spatial feature enhancement module together. And inputting the obtained fine-grained features into an edge feature sensing module, and finally fusing multi-scale features through a multi-scale jump connection technology to generate a mask pattern with higher discrimination. The method has the advantages that the network parameter quantity is reduced and key information is reserved through a feature selection mechanism; a spatial feature enhancement module is used for enhancing multi-scale feature representation and remote dependence modeling; the dilution of the semantic context is relieved by means of a semantic supplement module so as to improve the positioning precision; and an edge feature enhancement module is adopted to enhance edge semantic perception and improve boundary integrity. According to the method, the camouflage target detection performance is remarkably improved with relatively low calculation cost.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Construction state monitoring and risk assessment method and device based on BIM (Building Information Modeling) multi-mode conversion

The invention provides a construction state monitoring and risk assessment method and device based on BIM multi-mode conversion, and relates to the technical field of building information models. According to the method, a standardized image mode is generated by analyzing and extracting component information of a BIM model, and a BIM text mode is generated by using natural language description; constructing a graph structure mode based on space and construction logic, and realizing unified alignment and deep fusion of multi-modal data through multi-level modal alignment and a cross-modal attention mechanism to obtain a cross-modal fusion representation which is used for inputting a state recognition model and automatically detecting an execution deviation so as to monitor a construction state; and then introducing a deviation conduction mechanism to quantitatively calculate a comprehensive risk index of the component so as to carry out risk assessment. According to the method, the fusion representation which not only keeps semantic consistency but also conforms to construction logic can be obtained, the abstract cross-modal semantic features are converted into quantifiable and interpretable construction states and risk indexes, and powerful support is provided for intelligent analysis and application in a construction scene.
Owner:XIAMEN UNIV OF TECH

Watermarking method and system for large language model generated text

The invention relates to the field of watermarking algorithms, and provides a watermarking method and system for generating a text by a large language model. The method comprises the steps that a prompt text and a generated text sequence are input into a generation model, and an original logic score of a next mark position is generated through probability distribution calculation; performing semantic feature extraction on the generated text sequence by using an embedding model to obtain semantic embedding representation; converting the semantic embedding representation into a watermark logic score based on a pre-trained watermark model; and performing weighted combination on the original logic score and the watermark logic score to generate a final logic score, and outputting a watermark mark based on the final logic score. According to the method, the high robustness of the watermark is improved, and the safety of the watermark is also improved.
Owner:ZHONGJINKE INFORMATION TECH CO LTD +1

Semantic aerial view visual relocation method and device in non-exposed scene, electronic equipment, storage medium and program product

The invention provides a semantic aerial view visual relocation method and device in a non-exposed scene, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a multi-view image sequence under a non-exposed scene (such as a tunnel, an underground pipe gallery or an underground parking lot); semantic recognition is carried out based on a pre-trained semantic target detection model, and spatial consistency semantic features are extracted through a semantic-geometric dual-channel fusion mechanism combining a semantic mask and geometric constraints; the method comprises the following steps of: realizing three-dimensional reconstruction by using a voxel micro-renderable modeling method (VGGT), and generating a dense three-dimensional semantic point cloud fusing semantics and a geometric structure; two-dimensional semantics are mapped to a three-dimensional space through a projection and back projection relation, and point cloud semantics are endowed; main structure planes such as the ground, the left wall surface and the right wall surface are extracted, and a two-dimensional semantic aerial view with semantic annotation is generated; and pose estimation is carried out based on a reciprocal matching strategy guided by a semantic mask, so that visual repositioning with high precision, high robustness and semantic interpretability is realized. The method breaks through the problems of low precision, sparse features and poor semantic consistency of traditional visual repositioning in a non-exposed environment, and can be widely applied to the fields of intelligent transportation, underground inspection and unmanned system positioning.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Depth joint source channel coding method and related device

The deep joint source channel coding method comprises the following steps: a sending end obtains a first semantic feature of a first image, and performs constellation mapping modulation on the first semantic feature to obtain a first initial constellation point; the sending end extracts semantic information of the first initial constellation point, generates a rotation angle corresponding to the semantic information, and rotates the first initial constellation point according to the rotation angle to obtain a first rotation constellation point; the sending end generates a real part interleaving matrix and an imaginary part interleaving matrix based on the first rotating constellation point and the real part and the imaginary part of the channel state information, resorts the real part and the imaginary part of the symbol sequence respectively and then recombines the real part and the imaginary part into a complex signal to obtain a first interleaving signal; the sending end transmits the first interleaved signal to a receiving end through a fading channel, and the receiving end receives a second interleaved signal; the receiving end performs semantic reconstruction decoding on the second interleaved signal to obtain a second image; or, the receiving end performs semantic classification decoding on the second interlaced signal to obtain a second image category.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Path planning method and system for inspection robot

The invention belongs to the technical field of inspection robot systems, and particularly relates to an inspection robot path planning method and system.The visual semantic perception module collects an industrial environment image through a top industrial camera, after graying and Gaussian filtering preprocessing, feature points are detected and matched through an ORB algorithm, and a path planning result is obtained; in combination with an illumination self-adaptive threshold screening mechanism, mismatching points are eliminated, semantics are marked, and a semantic feature map is constructed; the initial path planning module generates an initial path through a semantic cost-containing A * algorithm based on the map; the dynamic obstacle avoidance module captures a moving obstacle by using a visual sensor, and predicts a trajectory through Kalman filtering; the path optimization module combines an initial path and an obstacle track, optimizes the path by using quadratic programming of a fusion curvature constraint and an energy consumption model, and corrects positioning by fusing vision and IMU data through a dynamic weight fusion algorithm; and the execution feedback module generates an instruction according to the optimized path, re-triggers path optimization, forms a closed loop, and ensures the inspection stability.
Owner:SICHUAN JOYOU DIGITAL TECH CO LTD

Hip joint osteophyte detection method based on texture perception and multi-scale feature adaptive fusion

The invention relates to the technical field of medical image analysis, computer vision and deep learning, in particular to a texture perception and multi-scale feature adaptive fusion hip joint osteophyte detection method. According to the method, firstly, a bone texture feature extraction module is used for carrying out feature extraction and enhancement on a hip joint X-ray image, a parallel double-branch structure is adopted, gradient and space structure features are extracted through a Sobel operator and pooling operation, and feature maps are fused; then, the fused features are input into a double-backbone network, the first backbone network extracts local fine textures and long-range dependence features by combining convolution, self-attention and a gating mechanism, and the second backbone network extracts multi-scale semantic features through hierarchical grouping and depth separable convolution; and double-trunk output is dynamically weighted and fused through an adaptive fusion module, features are optimized through a multi-scale semantic fusion module, and finally the position and confidence of osteophyte are output.
Owner:XIAN UNIV OF POSTS & TELECOMM

Small sample self-learning accurate identification method based on distillation knowledge migration

The invention discloses a small sample self-learning accurate identification method based on distillation knowledge migration. The method comprises the following steps: S1, extracting deep semantic features of a source domain and shallow features of a small number of samples of a target domain, and calculating a mapping matrix; s2, calculating an entropy difference distillation excitation function based on the initial alignment features; s3, executing domain knowledge distillation and generating staged distillation representation; s4, constructing a composite fitness function and initializing a parameter population; s5, performing iterative optimization by adopting a variable step size dynamic feedback compression strategy; s6, loading the optimal parameters and performing coupling alignment with the historical distillation representation; and S7, performing combined fine adjustment on the distillation weight and the model parameters through self-learning feedback. According to the method, through adaptive knowledge distillation and dynamic optimization feedback closed loop, high-precision and adaptive identification under extremely few labeled samples is realized, the generalization ability of the model is remarkably improved, and overfitting is effectively inhibited.
Owner:BEIJING KEANKE INTELLIGENT TECH CO LTD

Special equipment multi-mode interactive defect diagnosis system based on natural language processing

The invention relates to the technical field of natural language processing, in particular to a multi-modal interactive defect diagnosis system for special equipment based on natural language processing, which is characterized in that a feature extraction module obtains a voice instruction and a text report in the operation of the special equipment by constructing a multi-modal channel, extracts semantic features and the operation state of the equipment, and sends the semantic features and the operation state of the equipment to a database; the modal fusion module utilizes semantic association modeling to fuse features of different modalities to generate a heterogeneous semantic graph, and eliminates conflicts and redundancy between modalities through semantic slot consistency discrimination and redundancy compression coding optimization information; the defect diagnosis module is used for carrying out node representation learning by combining a graph neural network based on a heterogeneous semantic graph, identifying equipment defects and types thereof, and adding a multi-modal attention fusion layer; and the defect output module outputs feedback according to a diagnosis result to assist an operator in performing accurate defect diagnosis. The fault type of the special equipment is identified through the heterogeneous graph neural network.
Owner:FULIDA TECHNOLOGY CO LTD

Multi-mode re-identification method based on semantic-style decoupling distillation

The invention belongs to the technical field of image processing, tracking and recognition, and relates to a multi-mode re-recognition method based on semantic-style decoupling distillation. The method depends on a multi-modal re-identification model which comprises a multi-modal feature extractor comprising a teacher branch module and a student branch module, a decoupling distillation module and a hierarchical self-supervised learning module, and comprises the following steps: constructing a mixed multi-modal feature extractor sharing a shallow layer and an independent deep layer to extract mixed features; performing dual supervision of semantic distillation and style distillation, modeling modal-invariant semantic information and modal-specific style information, and realizing effective decoupling of a feature space; a hierarchical self-supervised learning space is constructed, and in combination with intra-modal and cross-modal comparative learning, images under local damage and style disturbance conditions are scrambled; according to the method, recognition performance and reasoning efficiency are both considered, semantic features and modal specificity styles are effectively separated, semantic consistency, feature robustness and network learning efficiency are cooperatively improved, and modal specificity is also reserved.
Owner:BEIJING INST OF TECH