Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

9results about How to "Accurate expression" patented technology

Image semantic communication method and system based on multi-modal large language model

The invention provides an image semantic communication method and system based on a multi-modal large language model (MLLMs), and aims to improve the semantic accuracy and system robustness of image transmission. The method comprises the following steps: extracting semantic features from an image source, detecting through an expert AI model, processing an OOD problem in combination with a general AI model, optimizing recognition precision by using a Bayesian method, encoding the features and transmitting the encoded features through a wireless channel. And after decoding by a receiving end, reconstructing the image by using a generation-evaluation framework based on the MLLMs and performing iterative optimization so as to ensure accurate reconstruction of semantic information of the image. According to the method, the OOD data processing capability is particularly enhanced, the image reconstruction quality is improved through the generation-evaluation framework, the limitation of a traditional model on new data distribution is solved, and the reliability of the system and the accuracy of image reconstruction are enhanced.
Owner:THE CHINESE UNIVERSITY OF HONG KONG

Robot interaction method and device, electronic device, and storage medium

This invention discloses a robot interaction method, device, electronic device, and storage medium. The method includes: detecting first pose parameters of the first facial target in response to a first condition being triggered by a first facial target; detecting first feature information of the first facial target in response to a angular difference between the facial orientation indicated by the first pose parameters and the optical axis of the target robot's camera being less than a preset angular difference, wherein the first feature information is used to indicate facial feature information and gesture feature information; determining the task intent associated with the first gesture feature information from multiple sets of task intents associated with the first gesture feature information based on the first gesture feature information in response to a matching degree greater than a preset matching degree of the first feature information and a third feature information among multiple sets of second feature information; and controlling the target robot to execute a corresponding feedback action based on the task intent associated with the first gesture feature information. This invention can improve the environmental anti-interference capability and intent recognition accuracy of robot interaction.
Owner:HANGZHOU ISOFTSTONE TIANQING ROBOT TECHNOLOGY CO LTD

Intelligent question and answer method and system for enterprise safety and environmental protection management

The invention discloses an intelligent question and answer method and system for enterprise safety and environmental protection management, and belongs to the technical field of intelligent question and answer. The method comprises the following steps: acquiring a security and environmental protection problem asked by a user, analyzing the security and environmental protection problem asked by the user by using a security and environmental protection semantic analysis model to obtain an analysis result, and judging whether the security and environmental protection problem asked by the current user belongs to the field of security and environmental protection systems or not according to the analysis result; if the method belongs to the field of security and environmental protection systems, rewriting a security and environmental protection problem proposed by a user by using a statement rewriting model to obtain an instruction statement; searching a rule unit in a security and environmental protection knowledge base based on the instruction statement to obtain a search result; and the recall summary model extracts content directly related to the user intention from the search result, and generates a safety and environmental protection answer. According to the method, the corresponding safety and environmental protection result can be generated in time according to the questions of the employees, the operation risk increase is reduced, and the safety and environmental protection system awareness rate of the employees is improved.
Owner:HAINAN PORT & SHIPPING INT PORT CO LTD

Gaussian Localization and Mapping System Based on Structure Awareness and Multi-Gaussian Refinement

PendingCN122550800ATracking stability advantageAdvantage map expression ability
This invention discloses a structure-aware Gaussian localization and mapping system based on multi-Gaussian refinement, belonging to the fields of computer vision and robot navigation technology. It aims to solve the problems of existing Gaussian localization and mapping systems, such as underutilization of scene geometry, easy Gaussian map degradation, unstable tracking, and poor map reconstruction quality. Taking RGB and depth images as input, it includes three core modules: structure-aware tracking, structure-enhanced mapping, and structure-driven keyframe management. It achieves color-independent stable tracking through joint modeling of global geometry and Gaussian covariance, employs a two-layer Gaussian map to balance real-time performance and reconstruction accuracy, introduces transparency regularization to reduce Gaussian redundancy, and combines structure-driven keyframe selection and dynamic weight allocation to optimize map performance. This invention improves the tracking robustness, map representation capability, and real-time performance of the system in complex environments, and can be widely applied in fields such as robot navigation and augmented reality.
Owner:BEIJING INST OF TECH

Image generation method and device, computer device and storage medium

PendingCN122597556AImprove production efficiencyImprove your grasp of details
The present disclosure provides an image generation method, device, computer equipment and storage medium. The method comprises: obtaining an initial image and an initial text; determining a segmentation scheme according to the size of the initial image, segmenting the initial image according to the segmentation scheme to generate at least one segmented image; using an image encoder to extract features of the initial image and the at least one segmented image respectively to obtain global features and local features, and fusing the global features and the local features to obtain an image encoding result; and inputting the image encoding result and the initial text into a preset image generation model to generate a target image.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Video search methods, devices, equipment, and storage media based on cross-modal retrieval

ActiveCN121117256BImprove the real-time performance of retrievalaccurate expressionMetadata video data retrievalSpecial data processing applicationsData transformationComputer graphics (images)
This invention provides a video search method, apparatus, device, and storage medium based on cross-modal retrieval, belonging to the field of data processing. The method includes: acquiring retrieval text data and converting it into background and foreground images; searching for each video to be searched based on the background images to obtain a first video search result; searching for each video to be searched based on the foreground images to obtain a second video search result; and determining the video search result corresponding to the retrieval text data based on the first and second video search results. This invention can accurately represent the background and target object of the retrieval text data, shorten retrieval time, and improve retrieval accuracy.
Owner:HEBEI XIONGXIN TECHNOLOGY CO LTD

Infrared and visible light vision information fusion method based on gradient transform prior

ActiveCN117173063BImprove fusion qualityaccurate expression
The application discloses an infrared and visible light visual information fusion method based on gradient transformation prior, and the method comprises the following steps: acquiring a registered infrared image and a registered visible light source image, wherein the registered infrared image and the registered visible light source image represent a consistent geometric transformation relationship in space; constructing a gradient transformation prior network; inputting the registered infrared image and the registered visible light source image into the gradient transformation prior network for feature extraction image fusion processing to obtain a reconstructed fusion image. By using the application, the detailed texture information in the visible light and the thermal radiation target information in the infrared image can be accurately extracted, and the fusion quality of the image is improved. The application can be widely applied to the field of image fusion technology as the infrared and visible light visual information fusion method based on the gradient transformation prior.
Owner:FOSHAN UNIVERSITY

Acoustic feature acquisition method and device, electronic equipment and readable storage medium

ActiveCN115662387Baccurate expressionexpress vividlySpeech synthesisEngineeringAcoustics
Embodiments of the present application provide an acoustic feature acquisition method and device and a readable storage medium. The method comprises: acquiring a target phoneme sequence; performing encoding processing on the target phoneme sequence based on a split self-attention mechanism to obtain a first feature matrix; and performing decoding processing on the first feature matrix based on the split self-attention mechanism to obtain an acoustic feature of the target phoneme sequence. The split self-attention mechanism converts the unified processing multi-head attention in the multi-head self-attention mechanism into split multi-head attention by performing scale transformation on the key-value matrix and the value matrix in a preprocessing manner, realizes modeling of information at different levels and different scales at the same time, takes into account the overall and detailed grasp, improves the feature analysis capability of the phoneme feature, better meets the needs of emotional speech synthesis, and makes the expression of the finally synthesized speech more accurate and vivid.
Owner:BEIJING SINOVOICE TECH CO LTD

A method of lemon health analysis

This invention discloses a lemon health analysis method, relating to the fields of intelligent agricultural monitoring and plant health diagnosis. It involves acquiring continuous images of lemon plants using a hyperspectral camera, and then using these images to detect changes in a dynamic differential spectral segmentation model. Sequence modeling is performed on the pixel-level multi-band spectral data. An adaptive motion-aware scheduler automatically suppresses loss when the proportion of motion change mask pixels exceeds a preset dizziness threshold to avoid background misjudgment, generating a pixel-level dynamic change map. This is compressed into a pathological dynamic state vector using a variational autoencoder, and end-to-end training is performed with action-conditional video prediction as the training objective. In the extrapolation phase, a recurrent dynamic model is used to perform multi-step imaginative prediction based on candidate action sequences, outputting the future evolution trajectory of the pathological dynamic state vector and generating a quantitative early warning or decision simulation report on the outbreak risk level. The system also includes change type classification and model closed-loop optimization functions.
Owner:WEISHAN JUFENG AGRI TECH CO LTD