Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

66 results about "Text mode" patented technology

Text mode is a computer display mode in which content is internally represented on a computer screen in terms of characters rather than individual pixels. Typically, the screen consists of a uniform rectangular grid of character cells, each of which contains one of the characters of a character set. Text mode is contrasted to all points addressable (APA) mode or other kinds of computer graphics modes.

Multi-modal pre-training model construction method and system for monitoring video

The invention provides a multi-modal pre-training model construction method and system for monitoring videos, and the method comprises the steps: automatically constructing a high-quality multi-modal alignment data set through a single-modal description generation model and a large-scale language model, and remarkably reducing the marking cost; a special coding network and a shared projection layer are adopted to realize feature extraction and uniform semantic space alignment of video, audio and text modes; performing dynamic semantic fusion by using a modal collaborative attention mechanism; designing cross-modal contrast learning, mask prediction and time sequence consistency tasks to carry out multi-task pre-training; an external knowledge base is introduced, and the semantic reasoning ability is enhanced through a microretrieval mechanism; and optimizing model parameters by adopting a multi-task joint loss function and an end-to-end training strategy. According to the method, efficient and automatic construction, deep semantic alignment and fusion and intelligent reasoning of knowledge enhancement of monitoring video multi-modal data are realized, and the understanding and generalization ability of the model in a complex scene is effectively improved.
Owner:BEIJING JIAOTONG UNIV

Hatred mold factor identification method and system based on dual particle size symmetry and context perception alignment

The invention discloses a hatred model factor identification method and system based on dual granularity symmetry and context perception alignment. The method comprises the following steps: firstly, respectively extracting a text mode and a visual mode from input image data containing text information, and carrying out feature coding and embedding so as to obtain text mode features and visual mode features; according to the method, a mode of constructing same-granularity and cross-granularity dual symmetric factors in visual and text modal representation is adopted to control modal internal feature fusion, so that a function of generating high-semantic-consistency modal internal feature representation is realized; in the inter-modal representation fusion process, a context perception cross-modal alignment strategy is introduced, and dual symmetric modal information is used as a prompt, so that the understanding of memetic context information is enhanced; whether aggressive or discrimination information aiming at individuals or groups is transmitted or not can be accurately identified by excavating potential emotional tendency and value orientation, and the identification accuracy and efficiency of the hatred model factor are improved.
Owner:NANJING AUDIT UNIV

Knowledge graph construction method and device, equipment, storage medium and computer program product

The invention provides a knowledge graph construction method and device, equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a script text and a deductive video corresponding to the script text; constructing a plurality of first nodes of a text mode based on the script text, and constructing a plurality of second nodes of an image mode and a plurality of third nodes of an audio mode based on the deductive video; based on the content information corresponding to the fourth node and the fifth node, determining a semantic consistency score between the fourth node and the fifth node; determining a time sequence synchronization score between the fourth node and the fifth node based on the time sequence information corresponding to the fourth node and the fifth node; in response to the fact that the semantic consistency score is larger than a first threshold value and the time sequence synchronism score is larger than a second threshold value, constructing a first relationship between the fourth node and the fifth node; and constructing the knowledge graph based on the plurality of first nodes, the plurality of second nodes, the plurality of third nodes and the plurality of first relationships.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Infrared and visible light image fusion method based on multi-semantic deep collaboration

The invention provides an infrared and visible light image fusion method based on multi-semantic deep collaboration. The method mainly solves the problem that an existing method cannot fully integrate text modals and global consistency between image fusion and downstream tasks. Comprising the following steps: 1) constructing a dual-task parallel network structure, and efficiently establishing a deep correlation between image fusion and a downstream segmentation task; 2) designing a multi-semantic deep collaboration module, realizing effective fusion of multi-modal information by deeply integrating text features, pixel-level features and segmented semantic features, and meeting semantic requirements of downstream tasks; 3) guiding an image fusion and segmentation task by using deeper and fine-grained semantic information in a text mode, and enhancing semantic consistency between a fusion result and a downstream task; and 4) inputting the obtained multiple semantic features into a fusion decoder to generate a final image fusion result. The semantic comprehension and visual perception capabilities of the model can be effectively enhanced, and the multi-modal image fusion performance is improved.
Owner:XIDIAN UNIV

Multi-modal federal learning method based on self-attention

The invention discloses a multi-modal federated learning method based on self-attention, and relates to the field of data processing, and the method comprises the steps: firstly, designing a multi-modal feature coding module at a client, carrying out the deep semantic modeling of image and text modals based on a Transform architecture, and capturing the internal structure information of each modality; secondly, a local cross-modal self-attention modeling module is introduced, fine-grained semantic alignment between images and texts is modeled through a bidirectional cross-attention mechanism, and the local cross-modal understanding ability is enhanced; a multi-head collaborative fusion optimization module is further combined, inter-modal deep semantic collaborative modeling is carried out under a multi-attention perspective, and a fusion effect is optimized through a gating mechanism, so that the problems of inter-modal conflicts and inconsistent expressions are effectively relieved.
Owner:ANHUI NORMAL UNIV

Cross-border consumption behavior dynamic analysis method and device based on large language model

The invention relates to the technical field of natural language processing, and discloses a cross-border consumption behavior dynamic analysis method and device based on a large language model.The method comprises the steps that multi-modal original data, obtained from multiple data sources, in a target cross-border consumption scene is obtained, the multi-modal original data at least comprises text data and visual data related to the target commodity or the target service; performing feature extraction and fusion processing on the multi-modal original data to generate a unified semantic representation containing text semantics and visual semantics; identifying a scene type to which the target cross-border consumption scene belongs, and determining contribution weights corresponding to a text mode and a visual mode in the unified semantic representation based on the scene type; and analyzing the weighted unified semantic representation according to the contribution weight, and generating a consumer behavior intention label corresponding to the target cross-border consumption scene. The cross-border consumption behavior intention analysis method and device can improve the suitability and accuracy of cross-border consumption behavior intention analysis, and meet the refined operation requirements of cross-border e-commerce.
Owner:SHENZHEN MINGXIN DIGITAL TECH CO LTD

Method, apparatus, device and medium for generating video in text mode

Methods, apparatuses, devices and media are provided for generating a video in a text mode in an information sharing application. In a method, a request is received for generating the video from a user of the information sharing application. An initial page is displayed for generating the video in the information sharing application, the initial page comprising an indication for entering a text. A text input is obtained from the user in response to a detection of a touch by the user in an area where the initial page locates. A video to be published in the information sharing application is generated based on the text input. In some examples, within the information sharing application, the user may directly generate a corresponding video based on a text input. In this way, a complexity of user operation may be reduced, and the user may be provided with richer publishing content.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Method of controlling autonomous vehicle and related apparatus

The invention discloses a method for controlling an autonomous vehicle and related equipment, and the method for controlling the autonomous vehicle comprises the steps: responding to a natural language instruction of a user, and generating a text feature based on the natural language instruction; collecting original sensor data, and inputting the original sensor data into a visual encoder of a VLA model for processing to obtain key features representing a current environment; the key features and the text features are input into the VLA model, a decision-making thinking chain output by the VLA model in a text mode is obtained, multimedia display is conducted on the decision-making thinking chain, and the decision-making thinking chain comprises the current state of a vehicle, decision-making reasons and a decision-making process.
Owner:SZ ZHUOYU TECH CO LTD

Document layout analysis and reconstruction method based on vector topology and conflict arbitration

The invention discloses a document layout analysis and reconstruction method based on vector topology and conflict arbitration, and relates to the technical field of document image processing, in particular to a method for carrying out accurate identification, classification and structured reconstruction on layout elements of an unstructured electronic document containing a complex vector chart. Constructing a page vector topological graph by using the bottom vector instruction and the connected topological structure; on the basis, multi-dimensional geometric features, statistical features and text mode features are introduced to jointly participate in conflict arbitration between the table and the complex graph; and applying a forced spatial exclusive constraint to text extraction by utilizing an arbitration result, and carrying out semantic classification and rearrangement in combination with a style reference and a spatial proximity relationship. According to the document layout analysis and reconstruction method, the distinguishing capacity of a table and a complex graph is improved, the attribution judgment precision of characters and a main body text in the graph is improved, and self-adaptive recognition of a title and a main body style is achieved under different layout styles.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Private domain community-oriented text and image multi-mode user interest identification method and system

The invention discloses a text and image multi-mode user interest identification method and system oriented to private domain communities. The system comprises a text feature extraction module, an image feature extraction module, a multi-modal fusion module, an interest probability aggregation module and a dynamic threshold judgment module. The multi-modal fusion module adopts a gating mechanism to realize self-adaptive fusion, a gating factor is automatically adjusted according to the characteristics of an input sample, and weighting is carried out between a text mode and an image mode, so that the robustness of a fusion result can still be ensured when the quality of information in different modes is unbalanced. And the interest probability aggregation module performs time sequence smoothing on the instant interest probability through an exponential weighted moving average method, and inhibits short-time noise interference. And the dynamic threshold value judgment module is used for calculating a dynamic threshold value according to the mean value and the standard deviation of the recent window, comparing the smoothed interest probability with the dynamic threshold value, and judging that the corresponding interest label is activated when the smoothed interest probability is greater than or equal to the dynamic threshold value. According to the method, the user interests can be accurately recognized in a complex and changeable private domain community environment, and the adaptability, stability and real-time performance of interest recognition are achieved.
Owner:北京娱广科技有限公司

Demonstration and understanding-oriented large-model multi-channel interactive joint output method

The invention discloses a demonstration and understanding-oriented large-model multi-channel interactive joint output method, which comprises the following steps of: S1, constructing a multi-channel interactive output model, performing information retrieval on a large language model LLM to obtain text data, and dividing the text data to obtain text fragment sequence data; s2, a non-text modal data generation module obtains non-text modal data through large language model (LLM) retrieval; a multi-modal rendering module performs multi-channel rendering on the non-text modal data to obtain multi-channel rendering data, and a synchronous controller performs synchronous association on the multi-channel rendering data and the text fragment sequence data based on the key position anchor point data; and S3, the multi-channel interactive output model performs synchronous associated output on the multi-channel rendering data and each text fragment. The invention provides a brand-new, more efficient and more attractive information interaction mode, and helps a user to more quickly and deeply understand the content output by a large model through the collaborative display of the text and the multi-modal information.
Owner:ZHEJIANG SHIZIZHIZI BIG DATA CO LTD

Text-prior multi-modal sentiment analysis method and system

The application provides a text-priority multimodal sentiment analysis method and system, which respectively constructs private expert networks of text, video and audio, is used for capturing unique sentiment expression features of each mode, simultaneously introduces a shared expert network to model general sentiment semantics across modes, so that collaborative modeling of multimodal information is realized, then a text mode is used as a leading mode to start a sentiment analysis process, a preliminary judgment is given based on text information, according to a confidence degree, video and audio information are gradually introduced according to needs to perform incremental supplement for sentiment classification prediction, a need-decoding strategy consistent with human cognition is realized, and a final sentiment category is obtained. Through simulating a human cognitive process of integrating multi-source information according to needs, the application reduces reasoning time and resource consumption, makes the model decision process more natural and interpretable, and is closer to the characteristics of different modal information being unbalanced and non-equivalent in actual application.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Interface anomaly detection method and electronic equipment

The invention relates to the technical field of software testing, and discloses an interface anomaly detection method and electronic equipment. The method comprises the following steps: respectively collecting a visual frame, a UI hierarchical tree, a log text and a resource index; respectively encoding the visual frame sequence, the UI hierarchical tree, the log text and the resource index through a multi-modal model to obtain encoding features; for any two modes in the visual mode, the structural mode, the text mode and the index mode, calculating attention weights between the two modes through a multi-mode model according to coding features corresponding to the two modes, and determining multi-mode representation after the two modes are fused according to the attention weights so as to obtain multi-mode representation of all pairwise mode combinations; determining a consistency energy value through a multi-modal model according to the multi-modal representation of all the pairwise modal combinations; and if the consistency energy value is greater than a first threshold value, determining that the to-be-tested software has interface abnormity. According to the method and the device, the interface abnormity can be accurately detected in real time.
Owner:AUTEL INTELLIGENT AUTOMOBILE CORP LTD

Video time retrieval method based on bidirectional semantic enhancement

A video time retrieval method based on bidirectional semantic enhancement comprises the following steps: firstly, respectively extracting video features and querying text features by using a pre-trained video encoder and a pre-trained text encoder; secondly, respectively inputting the extracted features into a feature alignment module to obtain aligned video features and text features; then, designing a TGVM module, dynamically enhancing video features according to global features of a text mode, and reducing irrelevant information in the video features; thirdly, designing a Feedback Decoder module and feeding back comparison loss, shortening the distance between decoded text features and original video features, and enhancing text feature representation through the original video features; then, using a cross attention mechanism to obtain joint features after interaction of the video and the text; finally, a corresponding time slice in the video is queried through a decoder prediction text, and a model is trained under the supervision of time retrieval joint loss; according to the method, the multi-modal reasoning capability of the multi-modal interaction part is enhanced, so that video and text information can be better aligned.
Owner:XIDIAN UNIV

An automatic driving three-dimensional scene data preprocessing method and system based on a large language model

The application discloses an automatic driving three-dimensional scene data preprocessing method and system based on a large language model, a text end generates a question paradigm for comparative learning based on a large language model for each category label, stimulates factual knowledge of the large language model, takes the factual knowledge as an answer space, generates a detailed category template for the category label of the automatic driving task, and caches the category template into an offline file for loading during training of a downstream model, expands the category template, and strengthens the most core category phrase; a visual end obtains key frames of an input video sequence through sparse sampling and dense sampling, uses a video random data enhancement method to perform image transformation on the obtained key frames, and enhances the robustness of a model to visual representation. The application processes information of a text mode and information of a visual mode respectively, fusion of different preprocessing methods can capture different prior knowledge, and complementary characteristics thereof are utilized to realize better performance.
Owner:JIANGSU UNIV

Method and system to display content from a PDF document on a small screen

Roughly described, a viewer application is provided for viewing a PDF document on a screen of a device such as a mobile phone or tablet. The viewer application may operate in page mode or in text mode. In page mode the original layout is maintained, and navigation assistance is provided by use of a navigation pane indicating the contents of the screen with a superimposed frame. Display of the navigation pane is controllable by the user. In page mode a selected text column is scrolled and zoomed to optimize reading. In text mode, text is extracted from the document and reformatted in text view to be continuous and complete in correct reading order, and images and advertising may be excluded. The user may toggle between page mode and text mode. The viewer application is implemented in software to by executed by a processor on the device.
Owner:BENDING SPOONS US INC

Voice representation model training method and device, equipment, storage medium and product

The invention discloses a voice representation model training method, device and equipment, a storage medium and a product, and the method comprises the steps: when a voice representation model is trained, carrying out the coding of obtained voice features, carrying out the token boundary prediction through employing the voice features, generating a token boundary, generating prediction text data according to the token boundary, carrying out the discretization of the prediction text data, and carrying out the training of a voice representation model. And discretized speech representation is obtained. Due to the fact that the token-level discretized speech features can be obtained, cross-modal expression with higher coupling degree of the speech mode and the text mode and closer information expression can be obtained, and semantic joint expression with more information is achieved, the coupling degree with a multi-modal language model is further improved, and the precision of a large model is prevented from being affected.
Owner:CHINA MOBILE COMM LTD RES INST +1

Multi-mode depression auxiliary detection algorithm and system

The invention discloses a multi-modal depression auxiliary detection algorithm, and the method comprises the steps: obtaining multi-modal data which comprises video data, voice data and scale data of a medical site; carrying out identity recognition, round labeling and format conversion processing on the multi-modal data to obtain structured data containing role information, time sequence round and content fields; performing feature extraction on the structured data to obtain a feature vector set containing an image mode, a text mode and a voice mode; carrying out fusion processing on the feature vector set by adopting a preset cross-modal attention mechanism dominated by a text mode to obtain multi-modal fusion feature information; and inputting the multi-modal fusion feature information into a pre-configured depression risk assessment model, performing feature transformation and compression through a full connection layer and a pooling layer, and outputting a depression disease probability for representing a target individual. According to the method, complementation and dynamic semantic modeling of multi-source data can be realized, and the accuracy and reliability of a detection result are improved.
Owner:JICAN ARTIFICIAL INTELLIGENCE LABORATORY (SHENZHEN) CO LTD +2

A multi-modal false information identification and credit punishment system and method

The application discloses a multi-modal false information recognition and credit punishment system and method, and particularly relates to the field of Internet information security, and comprises a multi-modal false information recognition module for receiving user-submitted information to be recognized, wherein the information to be recognized at least contains a text mode and an image mode; the module is integrated with a trained multi-modal neural network model.The multi-modal false information recognition and credit punishment system and method can extract text deep semantic features by using a pre-trained BERT model through a text processing submodule, extract image visual features by using a pre-trained VGG model through an image processing submodule, and perform weighted fusion on the text and image features through a feature fusion submodule based on an attention mechanism to mine deep inter-modal correlations, and finally output accurate recognition results through a classification output submodule, thereby effectively improving the accuracy and robustness of cross-modal false information recognition and coping with complex scenarios of text and image collaborative fraud.
Owner:DIGITAL JIANGMEN NETWORK CONSTR CO LTD

Data processing method and device, storage medium, equipment and program product

The invention discloses a data processing method and device, a storage medium, equipment and a program product, which are applied to a data processing scene. The method comprises the following steps: obtaining query data, wherein the query data comprises at least one modal data of a text modal, an image modal and an audio modal; based on the query data, a query feature vector of the target feature dimension is generated through a multi-modal retrieval model, and the multi-modal retrieval model is configured to be capable of outputting feature vectors of multiple different feature dimensions; and performing similarity calculation on the query feature vector of the target feature dimension and a candidate feature vector corresponding to each candidate data in the candidate data set to obtain a retrieval result. According to the method, by maintaining the multi-modal retrieval model, the output of the feature vectors of different feature dimensions is supported, the resource consumption is reduced, the flexibility is improved, and the retrieval performance can be ensured.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Malicious suffix construction method and controllable jail break attack method and device

The invention discloses a malicious suffix construction method, a controllable jailbreak attack method and a controllable jailbreak attack device, the method designs a jailbreak attack framework in a diffusion model SD, and the jailbreak attack method is operated from a text mode and an image mode. A transferable malicious suffix not containing sensitive words is learned in a text mode, a cosine loss function is used, the malicious suffix is close to malicious prompt semantics containing explicit sensitive words in an encoder layer, and after the learned malicious suffix is added to a jailbreak object prompted by an input text, the jailbreak object is prompted by the text. And converting the harmless prompt into a sensitive image generation instruction. And in an image mode, through a designed security checker loss and text consistent loss function, carrying out back propagation gradient optimization on the input original image to be edited, so that the generated image passes through the check of the security checker, and an image containing NSFW content is generated. According to the method, the security limitation of a security checker can be bypassed, the jail break attack is realized, the security vulnerabilities existing in the current text graph model are revealed, and the security protection measures of the model are evaluated.
Owner:CHINA UNIV OF MINING & TECH

User file updating method and system for large audio-visual model

The invention provides an audio-visual large model user file updating method and system. The method comprises the following steps: acquiring an audio-video stream in real time; extracting a plurality of face feature vectors and voiceprint feature vectors corresponding to the face feature vectors based on the audio and video streams; updating a pre-stored user file based on the face feature vector and the voiceprint feature vector to obtain a first user file; based on the audio and video streams and a preset sequence labeling model, obtaining effective dialogue data, and converting the effective dialogue data into text data; and updating the first user file based on the text data and the audio and video stream to obtain a second user file so as to update the user file. According to the audio-visual large model user file updating method and system provided by the invention, the pre-stored user file is updated through the face feature vector and the voiceprint feature vector, so that the user information is supplemented and perfected, the accuracy and timeliness of the user file data are effectively improved, and the user experience is improved. And the situation that the user file cannot be updated due to limitation to a text mode is avoided.
Owner:BEIJING XUANJI INTELLIGENT TECHNOLOGY CO LTD

Large model-based medical science popularization image-text generation method and system

The invention discloses a medical science popularization image-text generation method based on a large model, and belongs to the field of generative artificial intelligence. The method comprises the following steps: specifying a text description for generating a medical science popularization copywriting, and firstly analyzing, extracting and generating a required text instruction set and a picture instruction set; inputting the text instruction set into the natural language large model to generate related text contents, and orderly splicing the generated text contents to form a complete and coherent text part of the medical science popularization copywriting; meanwhile, the picture instruction set is input into the text graph large model, and a plurality of medical science popularization illustrations matched with the copywriting content are generated; performing content filling according to a preset medical science popularization copywriting typesetting template, and finally forming a complete medical science popularization copywriting finished product with image-text combination and clear structure. According to the method, full-process automation of medical science popularization content production is realized, the generation efficiency and professional accuracy are remarkably improved, and the readability and attraction of the content are enhanced in an image-text combination mode.
Owner:HANGZHOU HENGSHENG YUNTAI NETWORK TECH CO LTD

Method, system and device for image-text classification

PendingCN122368589AText modeRadiology
The application discloses a kind of image-text classification method, system and equipment, it is related to computer vision, natural language processing and other technical fields.Image-text classification method includes: original image data and text description data are respectively encoded and handled, and image feature data and text feature data are obtained;Image feature data and text feature data are spliced and handled, and spliced feature data is obtained;Spliced feature data is randomly discarded and handled, and image-text feature data is obtained;Image mode and text mode in image-text feature data are multi-modal complementary fusion processing, and target fusion feature data is obtained;Target fusion feature data is classified and handled based on multilayer perception, and image-text classification result is obtained.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

Active body discussion agent interaction method, system and device and medium

The invention discloses an active personal discussion agent interaction method, system and device and a medium, and the method comprises the steps: obtaining multi-mode perception information through an agent, extracting task-related information in the information, and converting the information into a unified text mode to obtain first information; performing state information updating on a preset multi-mode semantic knowledge graph according to the first information; evaluating the updated knowledge state based on the target state to obtain the current cognitive completeness of the intelligent agent; if the current cognitive completeness is smaller than a preset threshold value, extracting a semantic gap between the updated knowledge state and a target state; and converting the semantic gap into an action strategy based on the historical action track and the current space-time position of the intelligent agent, so that the intelligent agent interacts with the environment according to the action strategy to obtain new multi-modal perception information, repeating the steps until the current cognitive completeness is not less than a preset threshold value, and outputting a latest knowledge graph for display. The intelligent agent interaction efficiency can be improved.
Owner:GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY +1

Short message template generation method and device based on artificial intelligence, equipment and medium

The invention provides a short message template generation method and device based on artificial intelligence, equipment and a medium, and is applied to the field of intelligent medical treatment and finance, the method comprises the following steps: collecting short message behavior data of a user, and constructing a user portrait according to the short message behavior data; demand information input by the client is obtained, the demand information and the user portrait are input to a preset short message template generation model, a short message template of at least one mode is generated, and the short message template comprises a text mode, an image-text mode and a video mode; the short message template generation model comprises at least one of a large language model, a visual model and a multi-modal model; sending the short message template to a user, and obtaining feedback data of the user; and dynamically optimizing the short message template generation model according to the feedback data. According to the method, the short message template is generated through the model based on the demand information and the user portrait, and the model is optimized according to the feedback data, so that the generation efficiency, the updating rate and the personalization of the short message template are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent matching method of receipt file and voucher file and related device

The invention relates to an intelligent matching method of a receipt file and a voucher file and a related device. The method comprises the steps of obtaining a to-be-matched receipt file, and performing hierarchical analysis from a visual mode to a semantic mode through a text mode on the receipt file to obtain receipt data units corresponding to the receipt file; obtaining a to-be-matched voucher file, and performing structured analysis on the voucher file according to the table structure of the voucher file to obtain each voucher data unit corresponding to the voucher file; and matching each receipt data unit with each voucher data unit on a rule matching and vector matching double-layer surface, and integrating each pair of matched receipt data unit and voucher data unit to obtain an archived file. By adopting the method, the matching accuracy and processing efficiency in a complex format and semantic difference scene can be improved while manual participation is reduced, so that the requirements of high-frequency service processing and compliance tracing are met.
Owner:SHENGYE INFORMATION TECH SERVICE (SHENZHEN) CO LTD

Question reply method, apparatus and device, computer readable medium and program product

The embodiment of the invention discloses a question answering method, device and equipment, a computer readable medium and a program product. A specific embodiment of the method comprises the following steps: generating instance representation information and video feature representation information corresponding to each instance in a target video; generating instance prompt guide information according to the instance representation information and the video feature representation information; acquiring key feature representation information from the video feature representation information according to the instance prompt guide information; and inputting the key feature representation information and the text information corresponding to the target question into a large language model to obtain reply information corresponding to the target question. The implementation mode is related to artificial intelligence, and the key feature representation information can be accurately extracted from the target video on the basis of not losing instance-level information. On the basis, through the key feature representation information in the video mode and the text information in the text mode, the reply content can be accurately generated by utilizing a large language model.
Owner:JINGDONG CITY BEIJING DIGITS TECH CO LTD

Method for sensing track state based on data fusion of multi-mode monitoring of ballastless track

The present disclosure relates to the field of track monitoring data processing, and particularly relates to a method for fusing and perceiving track state based on multi-modal monitoring data of ballastless track, which is used to solve the problem that single detection or monitoring data is difficult to comprehensively and comprehensively reflect the actual state of track structure. The method comprises the following steps: acquiring multi-modal monitoring data, the multi-modal monitoring data is divided into text mode, visual mode and speech mode, obtaining text features about track structure deformation, track irregularity and structure vibration acceleration based on the text mode; obtaining visual features about structure apparent disease, interlayer separation, fastener loss based on the visual mode; obtaining audio features about vehicle-track structure vibration, track loosening abnormal sound and noise based on the speech mode; based on the text features, visual features and audio features, multi-modal fusion features are obtained to predict the track service state, which provides an effective means for accurate and comprehensive perception of structure service state.
Owner:BEIJING JIAOTONG UNIV +3