Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

437 results about "Code module" patented technology

The Code Module is a Networker-only Module that gives rewards and exclusive Items by entering a valid code. At the moment, only the code for the LEGO World Event Badge is known.

Foreign matter intelligent sorting robot control system based on AI recognition

The invention relates to the technical field of industrial robot control, and particularly discloses an intelligent foreign matter sorting robot control system based on AI recognition, which comprises a dynamic spatial feature extraction module, a manipulator motion state coding module, a collaborative conflict detection module, a dynamic trajectory optimization module and an execution control adjustment module, constructing a three-dimensional dynamic space model through multi-sensor fusion, and extracting spatial topological features by utilizing continuous coherence analysis; manipulator motion parameters are converted into topological space representation, and a track feature coding matrix is established; detecting interaction conflicts among the manipulators in real time by adopting a multi-scale coherence analysis method, and generating graded early warning signals; a collision avoidance track is optimized based on topological constraints and a virtual rejection field technology; precise execution is achieved through inverse kinematics of the Lie group theory and self-adaptive control.
Owner:SHANDONG JINING CANAL COAL MINE

Cross-modal AI-based traditional art gene decoding method and system

The invention provides a Chinese quintessence art gene decoding method and system based on cross-modal AI, and belongs to the crossing field of artificial intelligence and digital media art, and the method comprises the steps: S1, constructing a Chinese painting-music-text multi-modal data set; s2, inputting the traditional Chinese painting image into a visual encoder improved based on CLIP-ViT, and outputting a 512-dimensional visual Token sequence through a normalization module, a position encoding module and a Transform encoder; s3, inputting the visual Token sequence and the emotion label into a cross-modal adapter, and directly mapping the visual Token to a music hidden space by adopting a self-attention mechanism to obtain a music embedding vector; and S4, inputting the user parameters into the improved high-frequency fidelity generative adversarial network to generate Chinese traditional music audio conforming to five sound orders. According to the method, intelligent semantic communication between visual art and auditory art is realized.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Star flash mouse low-power-consumption communication system and method based on Polar code and SLE collaborative optimization

The invention belongs to the field of star flash mouse communication, and discloses a star flash mouse low-power-consumption communication system and method based on Polar code and SLE collaborative optimization, the system adopts a layered architecture design, collaborative optimization of energy efficiency and performance is realized through deep coupling of a physical layer and a link layer, and the performance of the system is improved. The core of the system is composed of three functional units including a Polar code coding module, an SLE control engine and a cross-layer optimization interface, and a complete closed-loop control system is formed. According to the invention, through deep cooperation of the Polar code and the SLE mechanism, the overall power consumption is reduced by 42.7% compared with the traditional scheme; according to the Polar code dynamic construction method based on track prediction, an ultra-low bit error rate of 10 <-5 > magnitude can still be kept in a 2.4 GHz complex interference environment, and in combination with signal-to-noise ratio self-adaptive coding parameter adjustment, it is ensured that transmission delay within 0.5 ms is always kept from a low-speed office scene to a high-speed electronic sports scene; in the aspect of engineering implementation, the method has the capability of quickly adapting to hardware platforms of different manufacturers, and a complete solution with high performance and low cost is provided for large-scale commercial use of the satellite flash technology in the consumer electronics field.
Owner:WUHAN PANSHENG DINGCHENG TECH CO LTD

Reference video object segmentation method and system based on motion modeling and multi-modal interaction

The invention discloses a reference video object segmentation method and system based on motion modeling and multi-modal interaction, and the method comprises the steps: taking a video sequence and natural language description as input, and generating a preliminary segmentation mask through a text coding and mask decoder; a Kalman filtering motion modeling module is introduced to predict the motion trail of the target object, and time sequence consistency optimization is carried out on the preliminary segmentation mask; fusing the historical track of the object and the action semantics in the semantic features on the basis of a key action semantic coding module to realize action semantic alignment and mask dynamic correction; the segmentation quality of the current frame is subjected to multi-dimensional scoring based on a representative frame screening mechanism, the representative frame is screened out to update a memory bank, and the long-term tracking stability is improved. According to the method, the problems of target drift, insufficient semantic alignment and memory pollution in a complex dynamic scene in the prior art are effectively solved, and the segmentation precision, robustness and semantic consistency are remarkably improved while the light weight of the model is kept.
Owner:ZHONGKE (SHENZHEN) WIRELESS SEMICON CO LTD

Three-dimensional scene reconstruction method and system based on monocular depth estimation

The invention discloses a three-dimensional scene reconstruction method and system based on monocular depth estimation, and belongs to the technical field of computer vision and three-dimensional reconstruction, and the method comprises the steps: carrying out the multi-scale feature coding of a monocular RGB image through a mixed attention depth coding module, and obtaining the hierarchical depth feature representation; carrying out autoregression depth decoding through a self-adaptive edge perception depth decoding module to generate an initial depth map; a depth confidence map is calculated through a geometric consistency constraint optimization module and is fed back to a coding module for iterative optimization, and a refined depth map is output; and three-dimensional Gaussian ellipsoid scene representation is constructed through the Gaussian ellipsoid scene reconstruction module. According to the invention, high-precision depth estimation and high-quality three-dimensional reconstruction are realized by constructing a depth-coupled closed-loop cooperative system.
Owner:HARBIN INST OF TECH

Birdsong classification method based on harmonic enhancement and time-frequency semantic joint modeling

The invention relates to the field of twitter recognition, in particular to a twitter classification method based on harmonic enhancement and time-frequency semantic joint modeling, which comprises the following steps: collecting twitter samples and carrying out noise reduction and standardized preprocessing, carrying out multi-scale convolution operation on Mel spectrograms by utilizing a layered acoustic encoder, extracting time-frequency features in combination with a channel attention mechanism, and classifying twitter classification results. The method comprises the following steps of: generating adaptive position codes through a dynamic time-frequency joint coding module, carrying out time-frequency mode modeling by combining a global-local interaction mechanism, introducing a semantic fusion module which comprises a frequency band pyramid unit, a harmonic enhancement unit and a time-frequency gating unit, realizing dynamic weighted fusion of multi-layer features, and carrying out time-frequency mode modeling through a global-local interaction mechanism. And inputting the fusion features into a classification layer, training a network by adopting a cross entropy loss function and a gradient descent algorithm, and outputting bird categories through a full connection layer, thereby solving the key problems of insufficient description of a non-stationary time-frequency mode, insufficient modeling of a harmonic structure, reduction of recognition performance in a complex noise environment and the like in the prior art.
Owner:HUNAN UNIV OF SCI & TECH

Transformer substation scene three-dimensional semantic segmentation method fusing local geometry and global context

The invention discloses a transformer substation scene three-dimensional semantic segmentation method fusing local geometry and global context, and the method comprises the steps: obtaining the three-dimensional point cloud data of a transformer substation scene, and constructing a semantic segmentation data set; a three-dimensional semantic segmentation model is constructed, and fusion features integrating local geometric information, global context information and neighborhood information are extracted through a down-sampling module, a position coding module, a local-global fusion module and a neighborhood propagation unit in sequence; integrating the multi-scale fusion features through an up-sampling propagation module, and outputting a semantic category prediction result by using a semantic segmentation head; designing a loss function, and training the model; and loading the trained three-dimensional semantic segmentation model to realize refined segmentation of the substation scene. According to the method, the understanding capability of the complex three-dimensional structure of the transformer substation is effectively enhanced, the segmentation precision of refined power equipment parts is remarkably improved, and reliable technical support is provided for intelligent operation and maintenance of the transformer substation.
Owner:ANHUI UNIV

Numerical control machine tool fault diagnosis system based on machine learning

The invention relates to the technical field of numerically-controlled machine tool diagnosis, and discloses a numerically-controlled machine tool fault diagnosis system based on machine learning. The system comprises a multi-source sensing data acquisition module for acquiring multi-dimensional sensing data such as vibration spectrum, spindle current waveform, temperature distribution, servo motor encoder feedback and the like; the operation feature coding module receives the multi-dimensional sensing data, extracts time domain statistical features and frequency domain energy distribution features, and generates a multi-source feature coding result; the incremental learning analysis module dynamically updates the feature weight through an incremental learning algorithm, and constructs an incremental training data set; the genetic optimization module optimizes the network structure and hyper-parameter configuration of the fault diagnosis model according to the incremental training data set, and generates optimized network structure parameters; and the integrated diagnosis decision module receives the current operation state data and the optimized network structure parameters, fuses diagnosis results of a plurality of base classifiers through an integrated learning algorithm, and outputs fault type classification signals.
Owner:DONGGUAN LONGCHENHUI MACHINERY EQUIPMENT CO LTD

Power load spatio-temporal dynamic knowledge graph construction and load prediction method

The invention relates to a power load spatio-temporal dynamic knowledge graph construction and load prediction method, which comprises the following steps of: constructing a text and digital sequence hybrid vector coding module, providing a hierarchical entity relationship joint extraction framework oriented to power system load data, constructing a Multi-Encoder-Bi-GRU-CRF power load entity recognition model, and constructing a power load entity model. Constructing a power load spatio-temporal dynamic knowledge graph in combination with a predefined relation rule base; meanwhile, time-space sub-graphs are divided, a space-time coupling self-adaptive adjacency matrix is constructed, and the space-time dependency relationship between nodes is quantified; and finally, combining the knowledge graph node embedded vector and the adjacency relation embedded vector, and jointly extracting the spatial feature and the time feature of the power load by adopting a space-time diagram convolutional neural network. Therefore, the load prediction algorithm provided by the invention not only can give full play to the advantages of multi-modal semantic integration and space-time modeling capability of the knowledge graph, but also can improve the load prediction precision, assist in realizing refined energy management of the power system and assist in making an optimal scheduling strategy, and has a good engineering application prospect.
Owner:TIANJIN UNIV +2

Small sample target detection system and method adaptive to airport complex scene

The invention relates to a small sample target detection system and method adaptive to an airport complex scene. The system comprises the steps that an image trunk feature extraction module extracts a multi-scale semantic feature map of an input image; the text embedding and coding module is used for coding and extracting an input text representing a category to obtain a semantic embedding vector; a category perception convolution kernel construction module extracts local visual features from the mesoscale feature map, and weights the local visual features to generate a category perception convolution kernel; a sliding convolution region matching module calculates the response intensity of each position and category perception convolution kernel in the multi-scale semantic feature map, and determines a center point based on the response intensity; the precise positioning module extracts a plurality of low-confidence threshold candidate frames to construct a candidate frame set, and selects the candidate frame closest to the center point from the candidate frame set as a target detection result; small target detection can be carried out by considering the detection speed, the positioning precision and the semantic generalization ability under the scene that samples are scarce and category features are easy to confuse.
Owner:WUHAN BRILLIANCE TECH CO LTD

Multi-modal image registration method and system based on deformation adaptation and computer equipment

The invention discloses a multi-modal image registration method and system based on deformation adaptation and computer equipment, and the method comprises the steps: collecting a plurality of groups of multi-modal images, carrying out the gray standardization, and constructing a diversified registration data set; building a registration network model comprising a pyramid coding module, a deformation adaptive module, a cross-modal interaction module and a registration parameter estimation module; inputting an image pair into the modules in sequence, respectively extracting basic feature mapping, deformation feature mapping and interaction enhancement feature mapping, and finally outputting an estimation conversion parameter matrix; a training process is supervised through a preset loss function, optimal network parameters are selected, and a trained registration model is obtained; in practical application, an image pair to be registered is input into the trained model, a conversion parameter matrix is obtained, and image registration is completed. The multi-modal image registration performance can be effectively improved, and the method still has good robustness and adaptability especially under the condition that serious geometric distortion and significant modal difference exist.
Owner:HUNAN UNIV

Interaction system and method with emotion dynamic evolution memory function, medium and processor

The invention relates to the technical field of artificial intelligence, in particular to an interaction system and method with an emotion dynamic evolution memory function, a medium and a processor. The interactive system comprises a sensing layer, a processing layer and a decision-making layer, the perception layer comprises a semantic text coding module and a fusion processing module; the semantic text coding module outputs the acquired audio as semantic features and high-dimensional features, and performs linear space conversion processing; the fusion processing module and the data output by the semantic text encoding module are linearly spliced to form multi-modal feature vectors, and the multi-modal feature vectors are classified into the current emotion state of the intelligent agent and the emotion state of the user; the processing layer is used for correcting the data output by the sensing layer; and the decision-making layer is used for carrying out information decision-making and storage on the data output by the processing layer. The technical problem that in the prior art, the number of labels is limited, undefined interaction modes are difficult to process, and consequently the character of a robot is limited is solved.
Owner:SOUNDLINK (NINGBO) INTELLIGENT TECHNOLOGY CO LTD

Mobile robot path optimization system and method supporting track mode switching

The invention belongs to the field of robot control, and particularly relates to a mobile robot path optimization system and method supporting track mode switching, and the system integrates the operation scene and the real-time load state of a mobile robot through a coding module, and generates a navigation task seed containing a path cost vector, a navigation strategy identifier and a scene conversion point parameter; and the central scheduler dynamically selects and calls a corresponding track navigation strategy sub-model or a flat ground navigation strategy sub-model according to the navigation task seed, and generates an executable track control instruction by combining with a space-time reservation table for realizing multi-vehicle conflict avoidance and resource scheduling based on a continuous value reservation weight. According to the method, load-adaptive path planning and multi-vehicle efficient cooperation are realized, smooth and seamless navigation switching between a track and a flat ground scene is guaranteed through the switching prediction sub-model, and the working efficiency and the system reliability in a mixed complex environment are remarkably improved.
Owner:SHANGHAI HENGZE FUHUI INTELLIGENT TECHNOLOGY CO LTD +1

Discovery platform for modernization of legacy program code

Methods and systems for improving modernization of legacy software using an intelligent discovery platform are described herein. A client-based agent may generate metadata regarding the received legacy software. The code metadata may be analyzed by a code classifier module, which computes a plurality of score factors from the metrics from the legacy software metadata using a knowledge base from a modernization platform. The classified code metadata may be used by a project-specific model to derive a plurality of sub-scores based on the plurality of score factors associated with the legacy software. An analytics engine may then identify a code module from the legacy software having a greatest derived vulnerability score factor. A graphical interface including reconstructed code, corresponding modern code, and an explanation of vulnerabilities may then be generated by the analytics and reporting component for the identified code module.
Owner:IONATE INC

Discovery platform for modernization of legacy program code

Methods and systems for improving modernization of legacy software using an intelligent discovery platform are described herein. A client-based agent may generate metadata regarding the received legacy software. The code metadata may be analyzed by a code classifier module, which computes a plurality of score factors from the metrics from the legacy software metadata using a knowledge base from a modernization platform. The classified code metadata may be used by a project-specific model to derive a plurality of sub-scores based on the plurality of score factors associated with the legacy software. An analytics engine may then identify a code module from the legacy software having a greatest derived vulnerability score factor. A graphical interface including reconstructed code, corresponding modern code, and an explanation of vulnerabilities may then be generated by the analytics and reporting component for the identified code module.
Owner:IONATE INC

Emotion recognition method and device based on multi-modal consensus and diversity decoupling

The invention relates to an emotion recognition method and device based on multi-modal consensus and diversity decoupling. The method comprises the following steps: firstly, collecting multi-modal input data including language, vision and audio signals and carrying out corresponding preprocessing; then, constructing a multi-modal consensus and diversity decoupling emotion recognition model which comprises a multi-modal decoupling coding module, a prototype-Gram unification module, a feature enhancement module, a diversity classification module and an emotion prediction head; then, inputting the preprocessed multi-modal input data into the multi-modal consensus and diversity decoupling emotion recognition model, and performing model training optimization based on a total loss function formed by emotion prediction task loss, decoupling loss, unified target loss and diversity loss; and finally, inputting the multi-modal data to be recognized into the trained multi-modal consensus and diversity decoupling emotion recognition model, and outputting an emotion recognition result. And the accuracy, robustness and interpretability of the multi-modal emotion recognition system are improved.
Owner:SICHUAN UNIV

Equipment invariance enhanced multi-modal deep learning model and application thereof

The invention discloses an equipment invariance enhanced multi-modal deep learning model and application thereof, and the model comprises an input module, a coding module, a modal fusion module, a classification module, an equipment adversarial branch module, an output module and a final total loss function. Extracting features through a coding module to obtain audio features and text features; after the audio features and the text features are processed by the modal fusion module to obtain final joint representation, the classification module processes the final joint representation to obtain classification results and corresponding probabilities, and the classification results and the corresponding probabilities are output by the output module; the equipment confrontation branch module enables an audio encoder to confront an equipment classification head in a training stage; and finally, introducing the total loss function in a training stage to optimize the model. The model provided by the invention has the capability of realizing high-accuracy identification on various respiratory system diseases on the premise of not depending on acquisition equipment of a specific brand, and shows excellent generalization performance and robustness in multi-equipment and multi-center data.
Owner:BOJIANG LIFE SCI (SHANGHAI) CO LTD

Cloud-side collaborative video content intelligent analysis and understanding system

The invention discloses a cloud-edge collaborative video content intelligent analysis and understanding system, which belongs to the technical field of video intelligent analysis, and comprises an edge intelligent sensing module, a semantic feature coding module, a cloud deep understanding module and a collaborative decision module, the coding module adopts Riemannian manifold mapping and quantum heuristic coding to realize efficient compression, the cloud module realizes deep semantic understanding through a graph neural network, the collaborative decision module constructs a closed-loop feedback mechanism to dynamically optimize parameters of each module, and the four core modules are deeply coupled to form a collaborative system. According to the method, cloud edge capability complementation, resource optimization configuration and continuous improvement of system performance are realized, the defects of a traditional scheme in the aspects of cloud edge cooperation capability, semantic understanding depth and adaptive optimization are effectively overcome, and an efficient, real-time and accurate technical scheme is provided for intelligent video analysis.
Owner:JIANGXI GAORUAN TECHNOLOGY CO LTD

Building material planning management system based on BIM technology

The invention discloses a building material planning management system based on a BIM technology, and relates to the technical field of BIM, and the system comprises the steps: a data collection module collects the full-process multi-source data of materials from purchasing to using; the unified coding module distributes a unique electronic identification code containing a category, a supplier and a full-process traceability index to each building material; the data storage and processing module cleans and converts the data and forms a standardized full-life-cycle database; the BIM integration module establishes one-to-one correspondence between material entities and BIM model components through electronic identification codes, and full-process data tracing calling and circulation state visualization are achieved; the interaction module provides a multi-terminal interface to support cooperation of multiple participants. According to the invention, the problems of information isolated island and traceability fault of traditional material management are solved, the management and control efficiency and collaboration are improved, and cost reduction and benefit increase of the construction project are facilitated.
Owner:ZHEJIANG ZHONGCHENG ZHINENG XINXI CO LTD

Cable partial discharge mode identification method and device based on multi-mode fusion, electronic equipment and storage medium

The invention discloses a cable partial discharge mode identification method and device based on multi-mode fusion, electronic equipment and a storage medium, and belongs to the technical field of cable partial discharge online monitoring. And converting the original signal into a multi-modal spectrum through phase resolution analysis, wavelet packet energy analysis, skewness and kurtosis statistics and recurrence plot analysis. The group of maps are input into a multi-modal fusion recognition model, the model analyzes the spatial distribution characteristics and the time sequence characteristics of the maps at the same time through parallel vision and time sequence coding modules, and two analysis results are subjected to weighted fusion, so that final classification recognition of the discharge mode is realized. Through the implementation of the method and the device, the problems that key information is lost due to dependence on single-dimensional feature representation and the identification accuracy is limited due to the fact that an existing identification model cannot analyze the spatial distribution characteristic and the time sequence evolution characteristic of the signal at the same time in the prior art can be solved.
Owner:ELECTRIC POWER RES INST OF GUANGDONG POWER GRID CO LTD

Response verbal skill generation method and device, electronic equipment and storage medium

The invention discloses a response verbal skill generation method and device, electronic equipment and a storage medium, relates to the technical field of computers, can be applied to financial science and technology and medical health business scenarios, and comprises the steps of randomly covering a single character for an original text of a user to generate a difference text; the pre-training intention recognition model comprises a text coding module, a feature fusion module, a multi-task collaboration layer and a classification layer, wherein the feature fusion module comprises a bidirectional long-short-term memory network and fuses an auto-encoder and an attention mechanism. The model is coded to obtain a vector containing overall and local semantics, an enhanced local vector is generated through a bidirectional long-short-term memory network, a global fusion vector is obtained through joint modeling, a multi-task layer optimizes slot positions to extract key information, and a classification layer combines the key information and the global vector to determine intention; and finally, calling the knowledge graph based on the intention, and generating a response verbal skill matched with the demand. According to the method, the response verbal skill matched with the real demand of the user can be accurately generated.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Target identification method based on time-space separation pulse Transform, electronic equipment and storage medium

The invention discloses a time-space separation pulse Transform-based target recognition method, electronic equipment and a storage medium. The method comprises the steps of obtaining event stream data which is output by a target neuromorphic visual sensor and is used for target recognition; inputting the event stream data into a pre-trained target pulse neural network, performing pulse coding on the event stream data through a coding module, and outputting pulse coding features; performing local space-time feature extraction on the pulse coding features through a feature extraction module, and outputting local space-time pulse features; performing global spatio-temporal feature extraction on the local spatio-temporal pulse features through a spatio-temporal separation pulse Transform module, and outputting global spatio-temporal enhancement features; and performing linear mapping on the global space-time enhancement features through a task output module, and outputting a target recognition result. According to the invention, while event-driven sparse calculation and low power consumption characteristics are maintained, the discrimination stability of weak, small, sparse and slow moving targets is significantly improved.
Owner:SUZHOU SHENZHITU TECHNOLOGY CO LTD

Code module decoupling method and device, equipment, medium and product

The invention discloses a code module decoupling method and device, a medium and a product, which are applied to the field of financial science and technology, and comprise the following steps: acquiring a source code, and constructing a baseline dependency graph based on the source code; obtaining an incremental code, generating an influence node set according to the incremental code, and generating a change influence subgraph according to the influence node set and the baseline dependency graph; loading a predefined architecture health model, and evaluating the change influence subgraph based on the architecture health model to determine a health score; and according to the health score, triggering a decoupling actuator of a corresponding level to perform decoupling. By constructing the baseline dependency graph, a basic reference is provided for subsequent analysis of a code dependency relationship, and the interference of different language grammar differences on analysis is eliminated. And only the changed part is analyzed, so that the low efficiency of full-dose scanning is avoided, and the analysis time consumption is shortened. According to the method, the architecture risk is quantified into a specific score, a clear triggering basis is provided for subsequent decoupling operation, dependence on subjective judgment is avoided, and hierarchical automatic decoupling is realized.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Graph enhanced double-memory collaborative knowledge tracking model based on ACT-R cognitive architecture

The invention relates to the technical field of knowledge tracking, and discloses a graph enhanced double-memory collaborative knowledge tracking model based on an ACT-R cognitive architecture. Comprising a static knowledge structure coding module based on hypergraph projection, a batch-level dynamic learning track construction and coding module, a cross-graph gating fusion mechanism, a sequence modeling module and an expert hybrid prediction module. According to the method, long-term stable structured semantic association between concepts in declarative memory is modeled through a static knowledge structure diagram, a dynamic learning trajectory diagram based on batch reconstruction is designed to accurately capture an evolution rule of a behavior sequence in programmed memory, and on the basis, a cross-diagram gating fusion mechanism and a hybrid expert mechanism are introduced, so that the evolution rule of the behavior sequence in the programmed memory is accurately captured. And self-adaptive fusion and multi-path decision of double-graph features are realized.
Owner:HARBIN NORMAL UNIVERSITY

High-speed fixed-point multiplication circuit

The invention discloses a high-speed fixed-point multiplication circuit. The multiplication circuit is mainly composed of a multiplier coding module, a partial product generation module, a partial product compression module and a traveling wave carry adder module. The multiplier coding module is composed of a radix-4-Booth coding algorithm and an opposite number generation module, and is used for carrying out three-bit block coding on input multiplication data and generating an opposite number of a multiplicand in advance. The partial product generation module generates a plurality of groups of partial products with symbol extension according to the multiplicand coded signal. And the partial product compression module adopts an improved Wallace compression structure to perform layered compression on the partial product. And the traveling wave carry adder module sums the two groups of results output by compression and outputs a multiplication result. According to the invention, multiplication accumulation series can be reduced, the switching times of invalid signals in the circuit can be reduced, and the longest delay path in an operation link can be shortened, so that fixed-point multiplication which is high in speed, low in power consumption and more favorable for a comprehensive tool in structure is realized.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Multimodal assembly action recognition method for comparing semantic query

The invention discloses a multi-modal assembly action recognition method for comparative semantic query, and relates to the technical field of man-machine cooperation assemblation.The method comprises the steps that a visual sensor is arranged on an assembly workbench to obtain an operator action video, a sampling frame sequence is obtained through random frame sampling, a skeleton sequence is obtained through human body posture estimation, and a skeleton sequence is obtained through human body posture estimation; and inputting an assembly action recognition model to complete recognition. The model comprises an image coding module, a skeleton coding module, a feature fusion module, a text coding module and a semantic comparison module which are used for extracting image and skeleton features, fusing features, coding preset category text description, comparing action features with category text features and outputting a result with the highest similarity, and a comparison loss function is adopted during training. According to the method, multi-modal information is fused, the problems of single-modal limitation and multi-modal semantic segmentation are solved, category text semantics are fully utilized, the fine-grained action recognition precision is improved, the over-fitting risk is reduced, and the generalization and task migration ability of the model in a dynamic industrial scene is enhanced.
Owner:ZHEJIANG UNIV

Intelligent clothing pattern generation system based on enhanced multi-modal generation

The invention relates to the technical field of intelligent costume design, and discloses an intelligent costume pattern generation system based on enhanced multi-modal generation, which comprises a multi-module input module, a layered attention fusion module, a self-adaptive position coding module and a hierarchical decoding module, the multi-module input module is used for receiving a clothing reference image, a version design demand text and structured data, and performing feature extraction respectively to obtain image features, semantic features and structural features; the layered attention fusion module adopts a three-layer attention mechanism to perform fusion processing on the image features, the text features and the structural features to obtain final fusion features; the self-adaptive coding module carries out self-adaptive coding on the final fusion feature and the model parameter of the clothing template, and outputs a coding result; and the hierarchical decoding module is used for decoding the coding result to generate a complete garment pattern.
Owner:SHANGHAI UNIV OF ENG SCI

Multi-modal signal identification and dialogue method based on large language model

A multi-modal signal recognition and dialogue method based on a large language model comprises the steps that a signal coding module based on time sequence modeling is constructed, preprocessing and feature extraction are conducted on input I / Q signal data, and semantic alignment pre-training of signal features and modulation type text description is achieved through a contrast learning mechanism; designing a multi-modal fusion architecture, mapping signal features to a hidden space of a large language model by adopting a signal projector, and realizing deep fusion of the signal features and text features through special markers; constructing a dialogue generation module based on a pre-trained large language model, receiving the fused multi-modal input, and generating a natural language answer about signal analysis; performing feature alignment by training a double-layer MLP projector to realize end-to-end multi-modal signal understanding and dialogue ability; an intelligent question-answering system in a reasoning stage is constructed, natural language interaction between a user and the system is realized through a predefined professional prompt word template and a signal feature fusion mechanism, and multi-dimensional signal analysis query requirements are supported. According to the invention, the organic combination of signal understanding and natural language generation is realized, and the accuracy of signal identification and the user interaction experience are improved.
Owner:ZHEJIANG UNIV OF TECH

Context understanding and memory management system and method in large-model multi-round dialogues

The invention discloses a context understanding and memory management system and method in large-model multi-round dialogues. The system comprises a five-layer distributed micro-service technology architecture; wherein the business logic layer is integrated with a hierarchical memory management module, a dynamic context coding module, a semantic association engine module and a personalized adaptation module which cooperate with one another; the hierarchical memory management module is responsible for retrieval, hierarchical storage and dynamic updating of dialogue history; the dynamic context coding module is responsible for generating structured context representation; the semantic association engine module is responsible for constructing a cross-round semantic association network; the personalized adaptation module is responsible for generating personalized candidate responses. The technical problems of context loss, memory attenuation, incoherent semantic understanding, insufficient personalized adaptation, low resource utilization rate and the like of an existing large-model multi-round dialogue system are solved, and logic consistency guarantee, memory durability maintenance and personalized experience improvement in a long dialogue scene are realized.
Owner:TRANSN IOL TECH CO LTD

Electric power overhaul video motion detection method based on multi-scale state space

The invention discloses an electric power overhaul video action detection method based on a multi-scale state space, and the method comprises the steps: firstly collecting long video data in an electric power overhaul process, segmenting the long video data into a plurality of time sequence segments, and carrying out the spatial feature coding and high-dimensional mapping of the plurality of time sequence segments, and obtaining a time sequence token sequence; then, an electric power overhaul video action detection network is constructed and trained, and the electric power overhaul video action detection network comprises a time sequence multi-scale coding module, a scale perception state fusion device and a multi-label prediction layer; and finally, performing action detection on a time sequence token sequence corresponding to the electric power overhaul long video data to be detected by adopting the trained electric power overhaul video action detection network to obtain an action category detection result. According to the method, timing sequence multi-scale coding, state space modeling and a scale perception feature fusion mechanism are fused, the short-time sudden action and the long-range dependency relationship can be captured at the same time, and the recognition capability and the positioning precision of the complex concurrent action in the electric power overhaul video are remarkably improved.
Owner:STATE GRID ANHUI ULTRA HIGH VOLTAGE CO +1