Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

81 results about "Language representation" patented technology

Representation is the production of meaning through language. Representation is an essential part of the process by which meaning is produced and exchanged between members of a culture. It does involve the use of language, of signs and images which stand for or represent things.

Web application internationalization

A system and method is described for internationalization of web pages by extracting translatable content from extensible mark-up language (XML), or similar data-centric meta-language representations of web pages or data used to build web pages. The extracted translatable content is stored in a translation task repository (TTR) accessible by the web developer and the translator. The XML representation is then modified to include selection control logic to select the appropriate translations for insertion into the final web page. The translator accesses the TTR to translate the appropriate content and saves the translations back to the TTR associated with the original translatable data. The translations are obtained from the TTR as selection cases for the selection control logic of the XML representation. As the XML is converted into the web source code, the selection logic and translations are embedded therein facilitating building the web site in multiple different languages.
Owner:ADOBE INC

Natural language generation

Techniques for determining when speech is directed at another individual of a dialog, and storing a representation of such user-directed speech for use as context when processing subsequently-received system-directed speech are described. A system receives audio data and / or video data and determines therefrom that speech in the audio data is user-directed. Based on this, the system determine whether the speech is able to be used to perform an action by the system. If the speech is able to be used to perform an action, the system stores a natural language representation of the speech. Thereafter, when the system receives system-directed speech, the system generates a rewrite of a natural language representation of the system-directed speech based on the previously-received user-directed speech. The system then determines output data responsive to the system-directed speech using the rewritten natural language representation.
Owner:AMAZON TECH INC

Retrieval augmentation system for unstructured tabular documents

A system for extracting a number of data elements from one or more data sources. A text representing tables using a markdown language is extracted from spreadsheets and or other grid-based documents. The text is provided to a language model with a prompt. The prompt may be a chain-of-thoughts prompt. The prompt includes several requests and / or steps that cause the language model to extract one or more tables from the spreadsheet and output the tables using the markdown language or a different markdown language. The tables extracted from the spreadsheet are converted into table chunks and indexed for retrieval by a retrieval augmented architecture. When a prompt to extract particular information from the spreadsheet is provided, one or more relevant table chunks are identified and provided to a language model for extraction. Using the language model to separate tables improves information extraction accuracy while maintaining downstream instructions.
Owner:AMERICAN INTERNATIONAL GROUP INC

Methods for real-time accent conversion and systems thereof

Techniques for real-time accent conversion are described herein. An example computing device receives an indication of a first accent and a second accent. The computing device further receives, via at least one microphone, speech content having the first accent. The computing device is configured to derive, using a first machine-learning algorithm trained with audio data including the first accent, a linguistic representation of the received speech content having the first accent. The computing device is configured to, based on the derived linguistic representation of the received speech content having the first accent, synthesize, using a second machine learning-algorithm trained with (i) audio data comprising the first accent and (ii) audio data including the second accent, audio data representative of the received speech content having the second accent. The computing device is configured to convert the synthesized audio data into a synthesized version of the received speech content having the second accent.
Owner:SANAS AI INC

Time sequence prediction method based on multi-modal contrast learning technology

According to the time series prediction method based on the multi-modal contrast learning technology, original multivariable time series data are converted into structured visual representation and language representation, multi-modal representation with consistent inner performance can be constructed without depending on external natural language or real image data, and the time series prediction method is high in practicability. The deep semantic understanding capability of the model on the complex operation state of the rail transit is effectively enhanced; a multi-modal contrast learning mechanism is introduced, visual and text modal representation is aligned in a shared embedding space, positive sample consistency is maximized through InfoNCE loss, negative sample interference is suppressed, and the robustness and generalization ability of time sequence features are remarkably improved; and the importance of each variable on a prediction task is dynamically evaluated by using the aligned multi-modal representation, and key variables are automatically screened, so that the redundant information interference is reduced, and the prediction precision and the calculation efficiency of the model in a high-dimensional multi-variable scene are also improved.
Owner:CRRC CHANGCHUN RAILWAY VEHICLES CO LTD

Urban park perception value evaluation method based on multi-modal data fusion

The invention discloses an urban park perception value evaluation method based on multi-modal data fusion, and the method comprises the steps: extracting a multi-dimensional perception tag and an emotion intensity score from park social media text comment data through a pre-trained language model, and carrying out the aggregation of the multi-dimensional perception tag and the emotion intensity score, and generating a subjective perception result; performing element identification on the image comment data of the park social media by using the three models to obtain different scene element coverage rates, character / facility counts and language representations of visual elements; constructing a plurality of objective environment indexes to describe different environment states, extracting variable information from three model output results, POI data, remote sensing land class proportion and road network information, and constructing mapping with the corresponding objective environment indexes to quantify each objective environment index; and aligning and outputting the subjective perception result and the objective environment index in the park and period dimensions. According to the method, automatic quantitative evaluation of subjective and objective characteristics of urban parks can be realized on the national scale, and standardized technical support is provided for urban renewal and ecological management.
Owner:NANJING INST OF GEOGRAPHY & LIMNOLOGY

Low-quality medical image classification method, system and device and medium

The invention discloses a low-quality medical image classification method, system and device and a medium, belongs to the technical field of artificial intelligence medicine, and aims to solve the technical problem of low accuracy of benign and malignant classification of low-quality medical images in the prior art. Comprising the steps of obtaining a sample image, performing probabilistic language conversion, constructing an nmODE network model, training the nmODE network model and performing real-time image classification. When probability language conversion is carried out, a probability language term set is constructed, and different brightness degrees of an image are described by using different probability languages in the probability language term set; carrying out probability language conversion on the obtained sample image by utilizing the probability language term set to obtain a single-channel probability language representation result of the sample image; and according to a plurality of single-channel probability language representation results of the sample image, obtaining a multi-channel feature map which can be input into the model. Through probabilistic language conversion, the language expression ability of low-quality medical image features can be enhanced, and the accuracy of low-quality medical image classification is improved.
Owner:SICHUAN UNIV

Cybersecurity event handling and enrichment system

A Cybersecurity Event Handling Processor (CEHP) and method for processing security alerts includes: a File System containing a Universal Target Schema (UTS) of target language representations (UTS JSONs); a Normalizer running Feature Extraction and Word Embeddings algorithms; a Tree Converter; and a Transformer running linguistic and structural matching algorithms. The CEHP: (a) captures threat events in one or more native formats generated by cybersecurity tools; (b) runs Feature Extraction and Word Embeddings algorithms for tokenization and categorization of the captured events to create normalized events; (c) converts the normalized events into trees and then translates the trees into event representations in JSON (or XML) format (Event JSONs); and (d) runs nearest neighbor and / or linguistic and structural matching algorithms to compare the Event JSONs to the UTS JSONs to generate output JSONs (Translation JSONs) from the UTS corresponding to the captured events.
Owner:NUHARBOR SECURITY INC

A small sample-based general image counting method and device

A small sample based general image counting method, comprising: performing feature extraction on a first image to obtain first features; performing attention calculation based on the first features, a general language representation and a general visual representation to obtain first example features of a first example target, wherein the general language representation is used to describe categories of different objects, the general visual representation is used to describe visual information of different objects, the first example target is an object related to an example frame in the first image, and the first example features include a first language representation and a first visual representation, the first language representation is used to describe a category of the first example target, and the first visual representation is used to describe visual information of the first example target; performing matching on the first example features and the first features to obtain a correlation feature map; and obtaining a first counting result related to the first example target based on the correlation feature map. The method can greatly improve the generalization of the counting algorithm.
Owner:HUAWEI TECH CO LTD

Translating programs from a first programming language to a second programming language

A program translation system may generate a first program source internal representation expressed using a source language representation of a source programming language and a first program target internal representation that is expressed using a target language representation of a target programming language. The system may interpret the first program source and target internal representations with first program data to generate a first program source and target results. The system may, when the source and target results are inconsistent, modify the target language representation. The system may, when the source and target results are consistent, generate a second program source internal representation of a second program. The system may use the second program source internal representation to generate a second program target internal representation of the second program. The system may use the second program target internal representation to generate a second program in the target programming language.
Owner:SCHLUMBERGER TECH CORP

A city park perception value evaluation method based on multi-modal data fusion

The application discloses a kind of urban park perception value evaluation methods based on multi-modal data fusion, utilize pre-trained language model to park social media text comment data extraction multidimensional perception label and sentiment intensity score, and subjective perception result is aggregated and generated;Park social media image comment data is identified using three kinds of models, obtain different scene element coverage, figure / facility count, and language representation of visual elements;A plurality of objective environment indexes are constructed to describe different environmental conditions, variable information is extracted from the output results of the three models, POI data, remote sensing land class proportion and road network information, and a corresponding objective environment index is constructed to map to quantify each objective environment index;Subjective perception result and objective environment index are aligned in park and period dimension and output.The present application can realize the automatic quantification evaluation of subjective and objective characteristics of urban park at national scale, and provide standardized technical support for urban renewal and ecological management.
Owner:NANJING INST OF GEOGRAPHY & LIMNOLOGY

Method for training a language representation model, method and apparatus for searching for a statement

Embodiments of the present application provide a method for training a language representation model, a method for searching for a statement, and an apparatus. The method includes: obtaining a target training statement, where the target training statement is obtained by collecting statements in a target field to which the language representation model is applied; training a pre-trained language representation model according to the target training statement to obtain a target language representation model, where the pre-trained language representation model sequentially includes a phrase feature extraction layer, a syntactic feature extraction layer, and a semantic feature extraction layer, and the input of some nodes in the i-th layer of the semantic feature extraction layer is the output of the j-th layer of the phrase feature extraction layer, and i and j are integers greater than or equal to 1. Through some embodiments of the present application, the running speed of the language representation model can be improved, and the parameters in the target language representation model can be made more suitable for application in the target field, thereby improving the accuracy of the language representation model.
Owner:阳光保险集团股份有限公司

A one-to-many multi-user semantic communication model and communication method

The present invention discloses a one-to-many multi-user semantic communication model and a communication method. The model is integrated by a sending end and a plurality of receiving ends that establish communication relationships with the sending end and are independent of each other. The method comprises the following steps: collecting text sentences of various types according to preset user needs; combining text sentences of various types into text sequences and converting them into digital ID sequences as sending information of the sending end; generating a communication signal for channel transmission at the sending end and sending the signal to each receiving end; each receiving end performs channel decoding and semantic decoding on the received communication signal to restore the original sentence sent by the sending end; and inputting the signal into a semantic recognizer based on a distilled bidirectional language representation pre-training model to output corresponding sentences according to user needs. Through the system model and communication method of the present invention, the transmission procedure of multi-user communication is simplified and the efficiency of information transmission is improved; and the receiving end is trained in combination with a transfer learning method to improve the training efficiency.
Owner:NANJING UNIV OF POSTS & TELECOMM

Pre-training Method for Language Representation Model Based on Generator-Discriminator Architecture

The present invention discloses a pre-training method for a language representation model based on a generator-discriminator architecture, including corpus data preprocessing, building a generator-discriminator architecture, pre-training using a span masking-replacement detection method, language representation model training, and model verification steps. The present invention solves the problem of inconsistent sample data between the pre-training method and downstream tasks. While increasing the difficulty of the pre-training method, it reduces the fine-tuning complexity of downstream tasks, so as to be able to generate a deep bidirectional language representation model and better apply to downstream tasks.
Owner:HEBEI NORMAL UNIV

A Text Semantic Retrieval Method in a Military Scenario

The present invention discloses a method for text semantic retrieval in a military scenario. First, based on a military pre-trained model, a dual semantic retrieval model is constructed, fine-tuned on a military semantic retrieval dataset to form a question-answer pair language representation model, and a semantic vector library of military text data is obtained offline. A secondary inverted index is constructed by means of vector clustering. Second, based on the military pre-trained model, a text retrieval and refinement model is constructed and fine-tuned on a military semantic retrieval and refinement dataset. In the face of a real-time retrieval task, the question sentence language representation model is used to obtain the semantic vector representation of the question sentence, retrieve through vector similarity calculation, recall a text set that meets the user's needs, and use the text retrieval and refinement model to accurately locate specific text data and feedback it to the user. This method can accurately locate the data required by the user in real time from a large amount of military text data and can be used in scenarios such as massive text search and retrieval-based question answering in a military scenario.
Owner:THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

Method and device for generating commodity semantic representation vector

The application discloses a commodity semantic representation vector generation method and device, and relates to the technical field of electronic commerce. A specific embodiment of the method comprises the following steps: performing word piece covering processing on each commodity title in a commodity title set respectively, obtaining a word piece index vector, a sample length vector and a text segment vector corresponding to each commodity title according to the commodity title after the word piece covering processing, and inputting the vectors into a pre-training model to obtain a first vector representation corresponding to each commodity title and a covering prediction result; then, adjusting parameters of the pre-training model to generate a semantic representation vector extraction model; and extracting vectors from the commodity title set by using the semantic representation vector extraction model to obtain a commodity semantic representation vector. The embodiment is more suitable for commodity language representation vector extraction, and can obtain a more accurate semantic representation vector, improves the semantic representation effect, and improves the use effect of the commodity semantic representation vector.
Owner:BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1

Language representation model system, pre-training method and apparatus, device, and medium

Disclosed are a language representation model system, a language representation model pre-training method, a natural language processing method, an electronic device, and a storage medium. The language representation model system includes: a word granularity language representation sub-model based on segmentation in units of words, and a phrase granularity language representation sub-model based on segmentation in units of words. The word granularity language representation sub-model is configured to output, based on a sentence segmented in units of words, a first semantic vector corresponding to a semantic expressed by each segmented word in the sentence. The phrase granularity language representation sub-model is configured to output, based on the sentence segmented in units of phrases, a second semantic vector corresponding to a semantic expressed by each segmented phrase in the sentence.
Owner:DOUYIN VISION CO LTD

Cross-lingual Text Representation Method, Apparatus, Device, and Storage Medium of Fusion Word Alignment Adapter Module

The present invention discloses a cross - language text representation method, apparatus, device and storage medium integrating a word alignment adapter module, which relates to technical fields such as artificial intelligence, natural language processing, text modeling, etc. The specific implementation solution is as follows: constructing a source - language - target - language parallel corpus dataset, and constructing a word alignment matrix for parallel sentences through an unsupervised word alignment algorithm; inserting a word alignment adapter between each sub - layer of a cross - language pre - trained model with a Transformer structure, and jointly training through masked language modeling and word alignment modeling to achieve semantic alignment of cross - language representation features; inputting the cross - language representation features generated by the word alignment adapter module into a task adapter, so as to implement various cross - language downstream tasks. It solves the problem in the prior art that the cross - language representation effect is poor due to the difficulty of forming a word alignment mapping for low - resource minority languages. According to the technology of the present application, the performance of cross - language text representation for low - resource minority languages and various cross - language downstream tasks is improved.
Owner:XINJIANG TECH INST OF PHYSICS & CHEM CHINESE ACAD OF SCI

Multilingual speech recognition models for speech processing systems and applications

The present disclosure relates to multilingual speech recognition models for speech processing systems and applications. In various examples, described herein are multilingual speech processing models for speech processing systems and applications. The systems and methods described herein can use an end-to-end model that is capable of performing both ASR processing and translation processing to generate text presented in various languages. For instance, a user can provide at least audio data representing speech and an indication of a target language for translating the speech. The model can then generate one or more audio representations associated with the speech and one or more language representations associated with the target language. Additionally, the model can combine the audio representations with the language representations (such as by performing stitching, masking, fusing, adding, etc.) to generate one or more combined representations. The model can then process the combined representations to generate text that corresponds to the speech and is presented in the target language.
Owner:NVIDIA CORP

Text main driving-based learner multi-modal sentiment analysis method and device

The application discloses a learner multi-modal sentiment analysis method and device based on text main driving, extracts multi-modal data with student related sentiment information embedded in an online classroom, uses a language representation pre-training model BERT to perform feature extraction on text modal data, uses an LSTM to perform feature extraction on audio and visual modal data, uses a cross-modal attention mechanism to fuse multi-modal information, and outputs a final result of multi-modal feature fusion, and performs sentiment analysis according to the final fusion result. The contrast learning technology is used to promote single-modal feature coding quality, maintain task related modal data uniqueness, ensure that multi-modal fusion results sufficiently learn unique sentiment information of various modal data generated in a classroom, improve learner participation in an online learning environment, and thus promote teaching quality.
Owner:ZHEJIANG NORMAL UNIV

Large model reasoning enhancement method combining knowledge graph embedding and gating residual connection

The invention discloses a large model reasoning enhancement method combining knowledge graph embedding and gating residual connection, and relates to the technical field of artificial intelligence and natural language process.The method comprises the steps that initial knowledge embedding is obtained through knowledge graph embedding learning, and a knowledge unit representation library is constructed through semantic space mapping processing; identifying a knowledge demand based on the text processing training data to obtain a query triple, retrieving a candidate knowledge set in the knowledge graph and constructing an enhanced training sample; inputting the enhanced sample into a large language model of an integrated knowledge gating residual connection module, and performing dynamic gating fusion processing according to the current layer hidden state and the knowledge unit representation to obtain an enhanced hidden state after knowledge injection; and generating a prediction result based on the enhanced hidden state, determining a joint loss function, and optimizing model parameters to obtain a reasoning enhanced large language model. According to the method, deep alignment and dynamic regulation and control of knowledge semantics and language representation are realized, and model reasoning accuracy and fact consistency are improved.
Owner:DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD

A novel word-level contrastive learning framework for sign language translation and a sign language translation system

The application discloses a novel word-level contrastive learning framework for sign language translation and a sign language translation system, relates to computer vision and sign linguistics, and provides a novel word-level contrastive learning framework ConSLT, which comprises a video input module, a visual extraction module, a sign language coding module, a sentence embedding module, a sign language decoding module, a contrastive learning module, a loss calculation module and an output module. The method comprises the following steps: 1) selecting a sign language corpus for modeling; 2) extracting sign language visual features; 3) performing end-to-end sign language video conversion; 4) calculating sentence embedding in a training stage; 5) constructing positive example pairs and negative example pairs; 6) calculating a sign language translation model loss; and 7) outputting a sign language translation result. From the perspective of natural language processing, the application explores contrastive learning of sign language translation, directly utilizes data itself as supervision information, and can learn good sign language representation under a low-resource condition, so that the sign language translation system is more accurate and fluent. The ConSLT framework is not limited by a model and is suitable for different models.
Owner:XIAMEN UNIV

Method and system for migrating web application firewall (WAF) configuration data across WAF providers

ActiveUS12695723B1Web applicationOrganizational context
The method and system for migrating web application firewall (WAF) rules across different WAF providers is presented. The method includes parsing a plurality of WAF rules from a plurality of WAF providers, wherein the plurality of WAF rules is expressed in varying provider-specific formats; enriching a source WAF rule of a source WAF provider with organizational context; generating, using a trained cross-provider semantic similarity model, a provider-agnostic language representation of the source WAF rule based on the provider-specific format of the source WAF rule and the organizational context; constructing a capability model of a target WAF provider; generating, using the trained cross-provider semantic similarity model, a target WAF rule for deployment in the target WAF provider based on the provider-agnostic language representation of the source WAF rule and compatible with the capability model; and coordinating a staged deployment of the target WAF rule in the target WAF provider.
Owner:HUSKEYS SECURITY LTD

Insurance operation process whole-link analysis method and device, equipment and medium

The application relates to the field of intelligent decision-making, and discloses an insurance operation process full-link analysis method and device, electronic equipment and a storage medium. The method comprises the following steps: defining a standard operation process of a service personnel; training a pre-trained hidden Markov model by using a marked text to obtain a trained hidden Markov model; training a pre-trained long short-term memory model by using a hidden state to obtain a trained long short-term memory model; inputting a small-stage distribution into a pre-trained language representation model to train the pre-trained language representation model by using the small-stage distribution, and obtaining a trained language representation model; analyzing a service operation process corresponding to a current interaction record by using the trained hidden Markov model, the trained long short-term memory model and the trained language representation model; and updating a question and answer knowledge base by using the service operation process to obtain an updated knowledge base. The application can finely manage an insurance operation process of a service personnel to a user.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Output text-oriented template code modification method and system based on deep learning

The present invention discloses a method and system for modifying output text-oriented template code based on deep learning, which relates to the field of computer software engineering technology. The method includes: generating a core language representation by semantically normalizing the template code to be modified, and completing the construction and restoration of a structured intermediate representation in a two-stage bidirectional framework; generating update instructions based on the user's editing operation on the output text, combining the metadata information of the structured intermediate representation, predicting and fusing the structured update instructions through a deep learning module, and finally restoring it to grammatically compliant template code through a backpropagation stage. The present invention realizes end-to-end mapping from output text editing to code updating by introducing a deep learning model, thereby improving the intelligence and interactivity of template code modification.
Owner:LONGYAN UNIV

Automatic low level operator loop generation, parallelization and vectorization for tensor computations

A method is provided for transforming a high-level language representation of a tensor computation graph into a low level language. The method includes assigning a tensor shape and a loop primitive. The method also includes generating, from the tensor computation graph and the assigned loop primitives, an initial loop structure. The method further includes positioning the layers of the tensor computation graph within a nested loop structure to provide a final loop structure, collapsing loops in the final loop structure, and mapping the collapsed loops to hardware components configured to execute the collapsed loops. The method can be applied to artificial intelligence (AI) and machine learning (ML) use cases for improved optimization of neural networks including compilation optimization for improving performance of simulations such as medical simulations, healthcare simulations, weather simulations, and / or simulations related to other complex systems, which can also support decision making.
Owner:NEC CORP

Retrieval method based on cross-modal semantics and hybrid counterfactual training

The present invention discloses a retrieval method based on cross-modal semantics and mixed counterfactual training, comprising the following steps: A. obtaining a reference image I R , target image I T and query text T Q The present invention addresses the following challenges: (1) modeling the feature representation of the image and text; (2) establishing a cross-modal representation modification module and a representation absorption and synthesis module to model visual language representations in a three-level cascade reasoning process; (3) constructing mixed counterfactual samples; (4) modeling global-local combinations to capture local-global information across different scales and modalities; and (5) deriving a final composite representation from the bottom-up hierarchical combination to capture implicit visual modifications and preservation in the reference image. (6) learning matching for combined retrieval of multiple modalities, then incorporating a θ-parameterized excitation. Finally, the learned composite image-text representation is uniquely aligned with the visual representation of the target ground-truth image. (7) evaluating the retrieval results using a loss function.
Owner:JIANGSU HUAZHEN INFORMATION TECH CO LTD

Entity question and answer generation method and device, equipment and storage medium

The embodiment of the invention provides an entity question and answer generation method and device, equipment and a storage medium. The method comprises the following steps: analyzing a to-be-processed knowledge file to obtain a processable text; according to the processable text, performing semantic segmentation by adopting a bidirectional language representation model to obtain one or more to-be-extracted text blocks; adopting a trained entity question and answer generation model to obtain a named entity corresponding to each to-be-extracted text block, and obtaining an entity question and answer pair corresponding to each named entity; the method comprises the following steps: acquiring a training named entity and a training question-answer pair for training an entity question-answer generation model through a cue word project and a language large model; and according to the question statement, screening out an entity question and answer pair matched with the question statement, and outputting an answer statement in the entity question and answer pair. The training named entities and the training question and answer pairs are generated through the cue word engineering and the language large model and are used for training the entity question and answer generation model, and the richness and the retrieval recall rate of a knowledge base are improved.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

A method for learning multi-view auxiliary representation and moment retrieval and related devices

This invention belongs to the field of computer vision and pattern recognition technology, and discloses a time retrieval method and related apparatus for learning multi-view auxiliary representations, aiming to solve the technical problem of unreliable time retrieval results in existing multimodal retrieval methods. The technical solution of this invention includes: constructing a multi-view auxiliary representation set based on text features; injecting target-specific visual attributes from the video into the auxiliary representation while maintaining semantic consistency to obtain an auxiliary representation that fuses visual attributes; achieving cross-modal association between the auxiliary representation and video features through multi-view semantic alignment to obtain enhanced visual features; and completing target time location through a detection head optimized by a loss function. The technical solution disclosed in this invention, through a multi-view auxiliary representation framework, actively injects visual evidence into the language representation, reduces the risk of overfitting sparse text cues, establishes semantic alignment of visual perception, and significantly improves the robustness and localization accuracy of time retrieval.
Owner:XI AN JIAOTONG UNIV

Data analysis method and system based on artificial intelligence

The invention discloses a data analysis method and system based on artificial intelligence, and relates to the technical field of artificial intelligence data process.The method comprises the steps that an input text is split, and semantic representation is output through a multi-scale pyramid; generating semantic anchor points through a dynamic time warping mechanism, aligning input semantic representations, constructing a domain knowledge structure, and converting the domain knowledge structure into a semantic prior tensor; inputting the semantic priori tensor into a priori injection channel of a text encoder, and encoding the semantic representation to form an encoded representation; constructing a semantic graph according to the coded representation and the semantic representation, inputting the semantic graph into a graph neural network, and outputting a graph enhanced representation; according to the coding representation and the graph enhancement representation, constructing a joint representation, training an analysis model, and outputting a cross-language representation; clustering the cross-language representation to generate an adversarial sample, and constructing a semantic causal path diagram with an original sample; the problems of semantic segmentation, alignment deviation and weak cross-language migration are solved.
Owner:SHANGHAI ZHIENTROPY INFORMATION TECH CO LTD