Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

23 results about "Input format" patented technology

Input formats are media formats that describe the basic properties of the media that you pass to the writer for encoding. For example, the frame size and color space of input video is described by the input format.

Active optical alignment method based on local lightweight language model

The invention discloses an active optical alignment method based on a local lightweight language model, and particularly relates to the technical field of precision installation and adjustment, comprising the following steps: extracting image indexes and platform attitude information through image acquisition, constructing a unified input format, verifying and inputting the language model; generating a six-degree-of-freedom pose adjustment strategy and a credibility score; the strategy is executed by the platform control device after being verified by a preset criterion; scoring is performed again after execution, and whether a feature posture is recorded or error compensation is triggered is judged according to a scoring result, so that subsequent input and strategy generation are optimized, and alignment precision and stability are improved; according to the method, the data consistency is guaranteed through a unified input format and a verification mechanism, the alignment precision is improved in combination with a dynamic scoring strategy, and the stability and fault tolerance of the system are enhanced by introducing a feature attitude recording and error back-filling mechanism, so that a strategy generation and execution process with high reliability, high adaptation and high robustness in an active optical alignment task is realized.
Owner:SHENZHEN CPT PRECISION TECH CO LTD

Systems and methods for generation of metadata by an artificial intelligence model based on context

Disclosed herein are methods and systems for generating metadata from content using one or more machine learning models. In an embodiment, a method may include receiving the content through a graphical user interface associated with the large language model, generating a first file by tokenizing the content into an input format for the large language model and merging the tokenized content with a content instruction, inputting the first file into the large language model, generating, using the large language model, metadata from at least the first file, the metadata reflecting a context associated with the content, generating a second file, the second file comprising the metadata, and displaying the generated metadata on the graphical user interface.
Owner:OPENAI OPCO LLC

Neural-network post-filter on value ranges and coding methods of syntax elements

A mechanism for processing video data is disclosed. The mechanism includes determining that a value of a neural-network post-filter characteristics (NNPFC) input format indicator (nnpfc_inp_format_inc) is coded as a u(N) coded syntax element where N is an integer greater than 0. A conversion is performed between a visual media data and a bitstream based on the NNPFC input format indicator.
Owner:BYTEDANCE INC

Friction stir welding defect detection method fused with weak supervised learning

The invention provides a friction stir welding defect detection method fused with weak supervised learning, which belongs to the technical field of welding defect detection, and constructs a unified detection framework suitable for various weak labels such as points, frames, graffiti and the like by referring to the advantages of a large model SAM (Section Anything Model) in visual segmentation. The method comprises the following steps: firstly, converting different types of weak tags into an input format acceptable by SAM by adopting a prompt adapter module, and enhancing the prompt compatibility of a model; and secondly, abnormal region responses are eliminated through a response filter, and the recognition precision of the disguise or low-contrast defect region is improved in combination with a semantic matcher. A prompt self-adaptive knowledge distillation mechanism is further introduced, so that knowledge migration from the SAM model to the lightweight detection model is realized, and the feature learning ability of complex weld defects is enhanced. According to the method, high-quality defect detection can be realized in the weld defect image without a large number of accurate labels, and particularly, the method has remarkable advantages in the aspect of processing the welding defects with fuzzy boundaries, small sizes or weak contrast, and has wide industrial application value.
Owner:GUANGXI UNIV

Value range of neural network post-processing filter about syntax element and encoding and decoding method

A mechanism for processing video data is disclosed. The mechanism includes determining that a value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfcinformatinc) is within a range from 0 to N (including a boundary value), where N is a positive integer. Conversion between the visual media data and the bitstream is performed based on the NNPFC input format indicator.
Owner:DOUYIN VISION CO LTD

Value range of neural network post-processing filter about syntax element and encoding and decoding method

A mechanism for processing video data is disclosed. The mechanism includes determining a value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfcinformatinc) to be encoded as a syntax element for u (N) encoding, where N is an integer greater than 0. Conversion between the visual media data and the bitstream is performed based on the NNPFC input format indicator.
Owner:DOUYIN CO LTD

Method and device for determining input length of question and answer model, equipment and medium

The invention discloses an input length determination method and device of a question and answer model, equipment and a medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: according to a preset long text of at least one language type, determining a candidate long text of at least one input format corresponding to the preset long text of each language type; performing text combination on the preset answer text and the candidate long text to obtain at least one candidate input text corresponding to the candidate long text of each input format; based on the target question and answer model, according to a preset question and answer instruction and the candidate input text, iteratively optimizing the text length of the candidate long text in the candidate input text to obtain at least one candidate input length of the target question and answer model; and determining a reliable input length of the target question and answer model according to the at least one candidate input length. According to the technical scheme, the reliable input length of the target question and answer model is explored through the combination mode of multiple answers and texts, and the reliability of determining the input length is improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

A patent text abstract extraction method, system, device and storage medium

The application discloses a patent text abstract extraction method, system and device and a storage medium. In the application, the patent text of a patent to be extracted is acquired; text embedding is performed on the patent text to obtain embedded text; the text data is converted into embedded text conforming to the input format of a subsequent model, so that the analysis efficiency of the model is improved; an abstract extraction model is used to perform abstract analysis on the embedded text to obtain abstract text; the abstract extraction model is obtained through combined cycle training of an expansion model, and the abstract model and the expansion model are both generated based on a BERT model; the expansion model is used to perform expansion analysis on the abstract text to obtain expansion text; through combined training of the abstract extraction model and the expansion model, the purpose of model training is strengthened, the accuracy of patent abstract extraction is effectively improved, and the application can be widely used in the technical field of text semantic processing.
Owner:ZHONGZHISHUTONG (BEIJING) INFORMATION TECH CO LTD

Lightweight lip language recognition model based on multi-teacher adaptive distillation

The invention relates to the technical field of lip language recognition, and discloses a lightweight lip language recognition model based on multi-teacher adaptive distillation, and the model comprises the steps: 1, carrying out the face detection, lip region-of-interest cutting and alignment of an input lip language video sequence, and obtaining a lip region-of-interest sequence; performing composite enhancement of directional design on the lip region-of-interest sequence; 2, extracting spatial-temporal features from the lip region-of-interest sequence by using a three-dimensional residual network variant to obtain unified sequence features adaptive to input formats of a multi-teacher model and a lightweight student model; and step 3, constructing a heterogeneous multi-teacher network. The technical scheme of heterogeneous multi-teacher network simultaneous prediction and weighted knowledge fusion is adopted, the technical effect of fully utilizing the complementary advantages of different teacher structures is achieved, knowledge breadth and depth synchronous migration is achieved, and the defects that the knowledge coverage range of a single teacher is limited and the generalization ability of a student model is insufficient are overcome.
Owner:CENT SOUTH UNIV

Sparse matrix feature extraction method based on multi-format input and parallel computing

The invention relates to a sparse matrix feature extraction method and system based on multi-format input and parallel computing, and the method comprises the steps: obtaining a to-be-processed sparse matrix from a sparse matrix file supporting multiple input formats, carrying out the format analysis, and obtaining global meta-information and a unified data view; according to the global meta-information and the unified data view, loading and dividing the sparse matrix in an MPI distributed environment to obtain distributed division information and a plurality of process local sub-blocks; according to the global meta-information and the distributed division information, non-zero elements of local sub-blocks of the processes are traversed in parallel by utilizing OpenMP multithreading in the local sub-blocks of the processes, local statistical features are determined, and a local imagination feature matrix with a fixed resolution is constructed; according to the global meta-information, reduction and aggregation are carried out on the local statistical features and the local imagination feature matrixes of the local sub-blocks of all processes through MPI communication, global unified global statistical features and global imagination feature matrixes are obtained, and the processing efficiency and format compatibility are improved.
Owner:HUNAN UNIV

Neural-network post-filter on value ranges and coding methods of syntax elements

A mechanism for processing video data is disclosed. The mechanism includes determining that a value of a neural-network post-filter characteristics (NNPPC) input format indicator (nnpfc_inp_format_inc) is in a range of 0 to N, inclusive, where N is a positive integer. A conversion is performed between a visual media data and a bitstream based on the NNPPC input format indicator.
Owner:DOUYIN VISION CO LTD

A large language model-oriented multi-modal benchmark evaluation method

PendingCN122285511ALinguistic modelEngineering
This invention discloses a multimodal benchmark evaluation method for large language models. First, structured information is extracted from the original PPT file, and a bimodal input format including structured JSON data and slide screenshots is constructed to ensure that the large model can simultaneously acquire structured layout information and complete visual information. A multi-level task module is constructed, including detection, understanding, modification, and generation tasks. For each of the constructed detection, understanding, modification, and generation tasks, corresponding evaluation indicators and calculation methods are designed to achieve reproducible evaluation of the large model's capabilities. This evaluation method can systematically characterize the differences in capabilities of large models in slide element recognition, layout reasoning, structured editing, and visual design generation, and can serve as an important benchmark for multimodal large models in the field of office document processing.
Owner:UNIV OF SCI & TECH OF CHINA

A Lossless Upscaling Method and System for Low-Resolution Images Based on a Large AI Model

This invention provides a method and system for lossless upscaling of low-resolution images based on a large AI model. The method includes: parallel image feature extraction from received images of different formats using a large AI model; determining a pixel generation strategy based on the feature extraction results, a pixel condition generation mechanism, and image content; performing resolution processing on the images based on the pixel generation strategy; filtering noise from the images based on the resolution processing results; automatically recommending an optimal upscaling ratio based on the noise filtering results and the image resolution; upscaling the images based on the optimal upscaling ratio; performing personalized adjustments based on user needs after upscaling; and exporting the personalized adjustments based on the input format. This method ensures that image quality is significantly maintained while upscaling low-resolution images, greatly improving image processing efficiency and guaranteeing lossless upscaling of low-resolution images.
Owner:SHENZHEN HAIHAI NETWORK TECH CO LTD

A method and apparatus for negotiation of conversational immersive audio session

Negotiation of a conversational immersive audio session includes obtaining one or more supported immersive conversational codec input formats, providing the one or more immersive conversational codec input formats in an order as a list, including an immersive conversational codec input format attribute in a session description file, populating the input format attribute with the list of immersive conversational codec input formats in the session description file, generating a session negotiation offer, based on the session description file, and transmitting the session negotiation offer to a receiver user equipment.
Owner:NOKIA TECHNOLOGIES OY

Combined Input Format Spatial Audio Encoding

PendingJP2026506746ASpeech analysisAudio frequencyInput format
1. An apparatus for encoding audio object parameters, the apparatus comprising: means for identifying one of at least two audio objects in an audio environment to be encoded separately; obtaining ratio parameters of the at least two audio objects for a time-frequency element of a frame comprising at least one time element and at least one frequency element; grouping ratio parameters associated with the at least two audio objects for the time-frequency element; quantizing the ratio parameters within the group, the quantization of the ratio parameters being configured to generate integer representations of ratio parameter values ​​that sum to a defined integer value for a particular group; encoding the integer representations of the ratio parameter values ​​in a first group as enumeration indexes; and encoding subsequent groups as difference values ​​relative to the first group.
Owner:NOKIA TECHNOLOGIES OY

Apparatus, method and computer program for encoding spatial audio content

Examples of the present disclosure relate to encoding spatial audio content using an immersive voice and audio service (IVAS) codec. In an example, an apparatus is configured to obtain a selected input format and encoding options for encoding spatial audio content, where the encoding options are configured to be switched to alternative encoding options for encoding the spatial audio content. The apparatus is also configured to receive an indication of an output format configured to be used by the playback device. The apparatus is further configured to upmix the selected input format to an alternative input format, where the alternative input format includes more channels than the selected input format, and the alternative input format is selected based at least in part on the output format.
Owner:NOKIA TECHNOLOGIES OY

Short video production method and system based on artificial intelligence

The application discloses a short video production method and system based on artificial intelligence, relates to the technical field of multimedia processing, and comprises the following steps: inputting a shot script into a multi-modal generation engine, generating a visual stream by adopting a lower triangular spatio-temporal attention matrix, and generating a text stream and an audio stream; calculating the neural conduction delay amount between the visual stream and the audio stream based on an audio-visual perception delay prediction model, and generating a corrected audio stream through time axis forward compensation; calculating a cross-modal attention weight matrix according to the node correlation strength of a dynamic causal diagram, projecting the text stream and the corrected audio stream to a joint feature space, and generating a multi-modal fusion feature; inputting the multi-modal fusion feature into a format adaptation engine, embedding a neural implicit watermark after spatio-temporal consistency verification, and encoding the neural implicit watermark into a short video. The application improves the semantic consistency and coordination between text, audio and visual content, thereby realizing high-quality and highly intelligent short video automatic generation as a whole.
Owner:BEIJING YIJIABANG TECH CO LTD

Sound wave hand account intelligent narrative generation engine

The invention relates to a sound wave hand account intelligent narrative generation engine, and relates to the technical field of multimedia content processing, the narrative quality is significantly improved: through multi-modal feature fusion and hierarchical narrative structure modeling, the generated narrative work is significantly superior to the traditional method in the aspects of structural integrity, emotional coherence and content attraction; the creation efficiency is greatly improved; the time and skill threshold required by manual editing are greatly reduced through automatic narrative generation; the intelligent interaction unit can dynamically adjust a narrative strategy according to user feedback, and can meet personalized requirements and preferences of different users in combination with diversified output options of the style adaptation module; the design of the method is compatible with an actual application environment, various input formats and output standards are supported, and the method can be seamlessly integrated into an existing content creation workflow.
Owner:SUPER SENSE DIGITAL TECHNOLOGY (DONGGUAN) CO LTD

An automated cross-project vulnerability verification method and apparatus based on a large language model

This invention provides an automated cross-project vulnerability verification method and apparatus based on a large language model, belonging to the field of vulnerability verification technology. It aims to solve the problem that existing technologies struggle to efficiently confirm whether upstream vulnerabilities can still be triggered in scenarios where there are significant differences in parameter semantics and input formats between the source and target projects. The method constructs a multi-agent parallel target-aware parameter transfer framework, combining RAG retrieval enhancement technology to automatically generate legal command-line parameter combinations related to the vulnerability function from the target software manual. Furthermore, it introduces the ReAct agent to embed the source PoC file into an input shell acceptable to the target software across formats without manual intervention, forming an initial cross-format seed. Subsequently, using function-level trajectory similarity as a feedback indicator, the parameter-seed pair is iteratively optimized, driving parameter-sensitive gray-box fuzzing to quickly converge to the precise input that can trigger the vulnerability.
Owner:XIAMEN UNIV OF TECH

Multi-modal task text processing optimization method based on large language model

The invention discloses a multi-modal task text processing optimization method based on a large language model, which relates to the technical field of computer text processing optimization, and comprises the following steps of: connecting and calling a locally or remotely deployed large language model, packaging text input into a request structure conforming to an input format of the large language model, by setting the large language model access end, the dynamic management optimization end and the processing optimization and feedback end, a perfect task text processing optimization system is constructed, a multi-modal input request is packaged and coordinated interaction is carried out, the text processing quality in a multi-modal task is remarkably improved, the multi-modal alignment error is reduced, and the task processing efficiency is improved. The context consistency and semantic integrity of the generated text are improved, the understanding ability and generalization ability of a large language model for a multi-modal scene are improved, text processing optimization data and corresponding analysis results are managed, visualized and stored, energy management can be achieved through Internet of Things cloud management and control, and the method is suitable for popularization and application. And the intelligent level of text processing optimization management is improved.
Owner:BEIJING HONGSHAN INFORMATION TECH RES CO LTD

Method of performing neural network filtering for video data and device

A device may be configured to perform filtering based on information included in a neural network post-filter characteristics message. In one example, the neural network post-filter characteristics message includes a syntax element indicating a purpose of a post-processing filter and a syntax element specifying whether syntax elements related to a purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristics message. The syntax element indicating a purpose may precede the syntax element specifying whether syntax elements are present.
Owner:SHARP KK

Systems and methods for converting documents with media into ai-ready and machine-readable formats

Systems and methods for converting documents with media into AI-ready and machine-readable formats are provided. A document conversion service is configured to convert a document from an input format to a machine-readable format by performing the operations of: inputting a document page of the document into a converter configured to convert text in the document page from the input format to the machine-readable format to yield a converted text page and to extract media data from the document to yield extracted media data. The operations also include inputting the extracted media data into the automated alt text service to generate extracted alt text and generating a media reference that comprises a path to the extracted media data. The operations also include inputting the extracted alt text and the media reference into a location in the converted text page where the corresponding extracted media data was extracted.
Owner:CABLE TELEVISION LAB INC

Apparatus, method and computer program for selecting modes of input format for audio stream

Examples of the present disclosure relate to selecting a mode of an input format for an audio stream comprising signals mixed from different sources. An indication of one or more modes available for the selected input format of the first audio signal is obtained. An indication of one or more modes available for the selected input format of the second audio signal is also obtained. The second audio signals are to be combined to form an audio stream. In an example, a mode is selected for an audio stream comprising both a first audio signal and a second audio signal, where the selection is based at least in part on one or more common modes available for a selected input format of the first audio signal and a selected input format of the second audio signal.
Owner:NOKIA TECHNOLOGIES OY