Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

258 results about "Language recognition" patented technology

Language Recognition. Summary: The goal of the NIST Language Recognition Evaluation (LRE) series is to establish the baseline of current performance capability for language recognition of conversational telephone speech and to lay the groundwork for further research efforts in the field.

Neck hanging type sign language interpretation equipment and sign language semantic recognition interpretation method thereof

The invention discloses a neck-hung sign language interpretation device and a sign language semantic recognition interpretation method thereof.The neck-hung sign language interpretation device comprises a wearing support which defines a neck surrounding space and comprises a middle connecting arm and two shells, and the two shells are connected to the two sides of the middle connecting arm respectively; the image acquisition unit is arranged on the front side of the shell, comprises a three-dimensional depth camera and an adjusting mechanism for driving the three-dimensional depth camera to rotate, and is used for dynamically tracking and acquiring continuous sign language videos; the processing module is integrated in the shell, comprises a main control processor and a memory, and is used for operating a sign language recognition algorithm; the voice output unit is electrically connected with the main control processor and is used for playing the translation result; and the power supply unit supplies power to the image acquisition unit, the processing module and the voice output unit. Gesture languages used by the deaf-mute for communication are converted into natural languages which can be easily understood by normal people, the deaf-mute group is helped to speak, and the problem of information output when the deaf-mute communicates with the normal people is solved.
Owner:SOUTHEAST UNIV

Sign language recognition glasses device, system and method

The invention discloses a sign language recognition glasses device, system and method, and belongs to the technical field of intelligent equipment, and the sign language recognition glasses device comprises a glasses frame which is used for supporting all components of the device; the camera is arranged on the front side of the glasses frame and is used for capturing hand actions in real time to generate image information; the data processing module is embedded in the glasses frame, electrically connected to the camera and used for processing the image information and executing sign language recognition; the main control unit is arranged on an ear rod of the glasses frame, is electrically connected to the data processing module and is used for coordinating the operation of each component; the display module is arranged at the positions of the lenses of the glasses frame, electrically connected to the main control unit and used for displaying character information obtained after sign language recognition, communication between hearing-impaired people and common people is not limited by specific places any more, and instantaneity and convenience of communication are greatly improved.
Owner:HANGZHOU YIDIAN ELECTRIC TECH CO LTD

Assembly line flow adaptive configuration method and system

The invention relates to the technical field of software engineering, and discloses an assembly line flow adaptive configuration method and system. The method comprises the following steps: obtaining source code data of a software project and performing static analysis and feature extraction to generate structured language description data; querying and matching in a preset template library based on the structured language description data to generate a task template set; performing task dependency relationship analysis and process topology assembly on the task template set to generate initial pipeline data; and performing dynamic strategy adjustment on the execution process of the initial pipeline data by adopting the real-time context data corresponding to the software project to generate target pipeline data. Through automatic language recognition and intelligent process generation, automatic configuration of the continuous integration process is realized, and the technical problem of low automation level caused by lack of automatic recognition and adaptation capability for a project development language environment in a project continuous integration and continuous deployment technology is solved.
Owner:SHANGHAI DONGFANG HOPE SOFTWARE TECH CO LTD

Automatic test script intelligent generation system

The invention discloses an automatic test script intelligent generation system, belongs to the technical field of software testing, and aims to solve the technical problems of how to realize automatic generation of test scripts, reduce dependence on professional technicians and improve test efficiency. The input module is used for inputting natural languages, extended prompt words, an initial browser and setting parameters; the intelligent analysis layer is used for realizing multi-language environment automatic test execution by combining the multi-language recognition capability of a large model and manual test cases of a set of languages; the element recognition layer is used for supporting a CSS / visual / semantic hybrid positioning strategy in combination with the visual recognition interface element capability and automatically selecting an optimal positioning scheme; and the execution feedback layer is used for acquiring latest page element information in real time when the UI structure is changed, and automatically utilizing the page semantic comprehension capability of the large model and the visual recognition capability of the multi-modal large model.
Owner:INSPUR QILU SOFTWARE IND

Sign language-voice conversion system

The invention discloses a sign language-voice conversion system, and belongs to the technical field of auxiliary communication and wearable computing. The system comprises a wearable myoelectricity acquisition module used for acquiring double-arm myoelectricity signals when a user executes sign language; the mobile terminal module is wirelessly connected with the acquisition module and is used for receiving and preprocessing the signal and uploading the signal; the cloud processing module is used for receiving the signal, converting the signal into text information through a sign language recognition model, and further calling a voice synthesis service to convert the text into voice data; and the wearable audio output module is used for receiving and playing the voice data. Through an innovative end-to-end hardware system architecture, natural, accurate and real-time translation and voice output of sign language gestures are realized, communication barriers between hearing-impaired people and healthy hearing people are effectively solved, and the system has the advantages of flexible deployment, user friendliness and privacy protection.
Owner:宋飞 +1

Multi-mode anti-interference communication method and system based on lip language recognition

The invention discloses a multi-mode anti-interference communication method and system based on lip language recognition, and belongs to the technical field of communication equipment. The method comprises the following steps: acquiring a face lip video stream and an audio signal; in response to the conventional mode trigger signal, feature extraction is performed on the lip video stream and the audio signal, and extraction results are fused to generate a fused feature vector; performing voice enhancement on the fusion feature vector in combination with lip motion information, and outputting an audio enhancement signal; and in response to the silent communication mode trigger signal, performing lip language recognition based on the face lip video stream to obtain a lip language recognition text, and converting the lip language recognition text into voice. Clear and stable communication in an ultra-strong noise environment can be realized by combining two types of modal information, and the problem that the communication quality is influenced by an existing high-noise environment and the requirement for mobile silent communication in a special scene are solved.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Full-language voice interaction legal affair agent system and method for small and micro enterprises

The invention discloses a full-language voice interaction legal affair agent system and method for small and micro enterprises. The core of the system comprises a full-language voice input module, a language recognition module, a voice recognition engine routing module, a special voice recognition engine cluster, an enterprise-level multi-language natural language processing module, an enterprise legal affair agent core, a multi-mode output module and a feedback optimization module, and all the modules work cooperatively. And seamless connection from voice input to scheme output is realized. The process of the method comprises voice acquisition and language recognition, audio routing and text conversion, semantic understanding and intention recognition, commercial risk multi-dimensional trade-off analysis, structured scheme generation and multi-modal output, and meanwhile, the system is driven to continuously iterate through a feedback optimization mechanism. According to the method, the working efficiency and experience of an enterprise owner are remarkably improved, the legal consultation revolution of'moving without using hands' is achieved, the method has low threshold and high specialty, legal instruments can be directly generated, and the method has high intelligence and self-adaptive capacity.
Owner:齐洪建

Sound environment analysis and monitoring method based on artificial intelligence

The invention discloses a sound environment analysis and monitoring method based on artificial intelligence, and belongs to the field of artificial intelligence. Sound feature extraction and parameter analysis; converting sound into text; converting the sound event into a text description form by using a multi-mode big language model with a sound event analysis function and outputting text information; timestamps are added to the output text information, and the text information is classified, sorted, recorded and stored for a user to trace back; determining a key sound event by identifying a dialogue scene; after the key sound event is triggered, the device reminds the user according to a specific form. Any equipment does not need to be implanted, and postoperative risks and maintenance cost do not exist; a user can perceive environment sound and understand dialogue content without listening by himself or herself, and a hearing aid is not needed; sign language actions do not need to be captured through a camera, and the problem of dialect sign language recognition errors does not exist; the user does not need to stare at the character display device in real time during use, and the user is reminded in a specific mode after the key sound event is triggered.
Owner:SHENZHEN TECH UNIV

Sign language recognition method based on double-arm electromyographic signals

The invention discloses a sign language recognition method based on double-arm electromyographic signals, and belongs to the technical field of human-computer interaction and biological signal processing. The method comprises the steps that multi-channel electromyographic signals generated when sign language gestures are executed are synchronously collected through electromyographic arm rings worn on the left forearm and the right forearm of a user; performing preprocessing and feature extraction on the signal to obtain a time sequence feature sequence of left and right arms; the time sequence feature sequence is input into a pre-trained two-arm collaborative recognition model, and the model outputs a sign language gesture recognition result by fusing the spatial-temporal features of the left arm and the right arm; and finally, the recognized text information is converted into voice to be output. According to the method, the cooperation and time sequence relation of the double-arm electromyographic signals is creatively utilized, the problems that a traditional visual recognition method is greatly interfered by the environment, privacy is invaded, and double-hand linkage complex gestures cannot be effectively analyzed through single-arm electromyographic recognition are solved, and natural, accurate and real-time recognition and translation of the double-hand sign language gestures are achieved.
Owner:宋飞 +1

Multi-language intelligent analysis system for medical documents

The invention provides a medical document multi-language intelligent analysis system, relates to the field of language translation, and improves the accuracy and efficiency of professional term translation. The method comprises the following steps of: firstly, performing language recognition on an original text document by a recognition module through text unitization and context vector generation, and matching a corresponding corpus; then, a translation module carries out lexical element alignment on the source lexical elements through term bank injection and an AI model, the translation process is automatically optimized, and accurate translation of the terminologies is ensured; and finally, the reconstruction module accurately replaces corresponding contents in the original text document with translation output through the mapping file, so as to ensure that the document format and typesetting are consistent. Through the automatic and optimized translation process, the quality and efficiency of professional term translation are remarkably improved, manual intervention is reduced, and the translation requirement of a high professional standard is met.
Owner:LUNAN PHARMA GROUP CORPORATION +2

Intelligent speech translation mobile phone and system capable of realizing multi-language inter-translation

The invention belongs to the technical field of mobile terminal communication and language translation, and particularly relates to an intelligent speech translation mobile phone and system capable of realizing multilingual inter-translation, which are characterized in that a two-way speech separation module and unit, a language recognition related module and unit, an end-cloud collaborative translation module and unit and a system are compatible with related module and unit collaboration; two-way voice collection, separation, language recognition and end-cloud collaborative translation are completed in a call scene, the system is compatible with a mainstream mobile operating system and communication application, a resource isolation and process linkage mechanism is adopted to guarantee operation compatibility, multilingual communication obstacles in the call scene are effectively solved, complex environment interference is adapted, the real-time performance and accuracy of translation are considered, and the system is suitable for large-scale popularization and application. And the native function of the terminal does not need to be modified.
Owner:SHENZHEN GUO ELECTRONIC INFORMATION CO LTD

Cross-border logistics single-multi-language machine translation method based on natural language processing

The invention discloses a cross-border logistics document multi-language machine translation method based on natural language processing, and relates to the technical field of machine translation, and the method comprises the steps: collecting original data of a cross-border logistics document, and carrying out text extraction and preprocessing through an OCR technology; performing language recognition and format analysis on the preprocessed document text to determine a source language and a target language; using a pre-trained cross-language alignment translation model to translate the document text into a target language text; the translation result is input into a compliance auditing module, and automatic compliance auditing is conducted through a knowledge graph and a double-check algorithm; and generating an output containing a translation result and a compliance audit report, and supporting manual review and revision. According to the method, the translation accuracy of the documents can be effectively improved, the manual auditing burden is reduced, and the globalization requirements of cross-border e-commerce and logistics enterprises are met.
Owner:QINGDAO UNIV OF TECH +1

Application memory overflow fault positioning method, system and device and medium

The invention discloses an application memory overflow fault positioning method, system and device and a medium, and the method comprises the steps: configuring a detection strategy in an application container node, carrying out the snapshot capture processing of a memory overflow event according to the detection strategy to obtain a memory snapshot, carrying out the application language recognition processing of the memory snapshot to obtain an application language, and storing the application language. And performing fault mode identification processing on the memory snapshot according to the application language to obtain a fault mode, performing analysis processing on the memory use condition of the application container node according to the fault mode to obtain an analysis result, and performing memory tracking processing on the analysis result to obtain a fault positioning result. The embodiment of the invention can automatically process the memory overflow fault, and can be widely applied to the technical field of computers.
Owner:SHENZHEN LANYOU TECHNOLOGY CO LTD

Continuous sign language recognition method based on context awareness convolution and subnet regularization

The invention discloses a context perception convolution and sub-network regularization-based continuous sign language recognition method, which comprises the following steps of: recognizing a sign language video by adopting a trained recognition network model, and inputting the sign language video to be recognized into a visual feature extractor; the features extracted in each stage are sequentially processed through a dynamic context sensing convolution module, residual connection is carried out on the features and original features, and final features of the current stage are obtained and transmitted to the next stage; and inputting the final output of the visual feature extractor into a time feature extractor, further modeling a time sequence dependency relationship, and obtaining an identification result through a classifier. According to the method, a dynamic context awareness convolution mechanism and a sub-network regularization method are introduced, so that the modeling capability of the model for spatio-temporal dynamic characteristics is remarkably enhanced, and the overfitting problem caused by a CTC peak phenomenon is effectively relieved.
Owner:ZHEJIANG UNIV OF TECH

Sign language recognition training data generation method and system based on multi-modal data enhancement

The invention relates to the technical field of graphic data reading, and particularly discloses a sign language recognition training data generation method and system based on multi-modal data enhancement, and the method comprises the steps: a sign language recognition assembly reads multi-modal data of a to-be-recognized sample based on a sampling device, and achieves multi-modal synchronous resampling; performing graph quality evaluation frame by frame after resampling, and outputting a graph quality evaluation result; screening the boundary candidate frame set frame by frame based on a graph quality evaluation result to form an action slice group of the to-be-identified sample; and finally, firstly matching a data enhancement initial strategy for each action slice, then optimizing the initial strategy based on a dynamic and static hierarchical labeling result to obtain a data enhancement execution strategy, executing multi-modal consistent data enhancement on the action slices according to the strategy, and writing the data enhancement into a training sample set. And the enhancement type and strength are adaptively adjusted along with the slice quality form and the dynamic and static structure, and finally a training data set for sign language recognition component training is formed.
Owner:SHANDONG OPEN UNIV +1

Sign language recognition model construction method based on uncertainty sample screening and comparative learning

The invention discloses a sign language recognition model construction method based on uncertainty sample screening and comparative learning, and the method comprises the following steps: S1, obtaining effective motion frames of each sign language vocabulary performed by each sign language performer, constructing an effective motion frame sequence, labeling the effective motion frame sequence to form a data set, and dividing the data set into a training set, a verification set and a test set; s2, inputting the training set into a sign language vocabulary label prediction model for prediction to obtain the sampling probability of each sample; s3, performing sampling probability estimation of a time stage on each sample in the training set to obtain the sampling probability of each sample in the time stage; s4, based on the sampling probability of each sample in the sampling probability time stage based on the sample loss, screening out high-value samples; and S5, constructing a basic recognition model, training the basic recognition model based on calculation results of the steps S2, S3 and S4 by adopting samples in the training set, and obtaining a sign language recognition model.
Owner:HEFEI UNIV OF TECH

Voice-to-text optimization method based on Ai assistance

The invention relates to the technical field of language recognition, and discloses a voice-to-text optimization method based on Ai assistance, which realizes global collaborative optimization through multi-scale feature adaptive extraction, joint modeling and a continuous learning mechanism. Compared with the problem of misrecognition amplification caused by a modular processing flow in the prior art, the method has the advantage that the influence of environmental noise on the upstream module is reduced by dynamically adjusting the feature extraction parameters and performing joint coding. Meanwhile, by introducing domain knowledge adaptation and context memory mechanisms, the defects of a traditional method in long-distance semantic dependency capture and personalized adaptive capacity are overcome, and therefore the overall performance of a voice-to-text system is improved.
Owner:HARBIN UNIV

Language identification method, device and equipment

The invention provides a language recognition method, device and equipment, and is applied to the field of voice processing. The language recognition method comprises the following steps: acquiring voice data; language recognition is carried out on the voice data to obtain a recognition result, and the recognition result comprises the first language; converting the voice data into text data in a first language to obtain first text data; performing quality inspection on the first text data to obtain a first inspection result; and under the condition that the first check result meets the text quality requirement, determining that the target language of the voice data is the first language. Therefore, by converting the voice data into the text data under the recognition language and performing quality inspection on the text data, the accuracy of language recognition on the voice data is improved.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

Method, system and device for realizing configuration change detection processing based on large language model, processor and storage medium thereof

The invention relates to a method for realizing configuration change detection processing based on a large language model, which comprises the following steps: acquiring a configuration change execution scheme, and outputting configuration change key information by utilizing the large language model; obtaining current configuration information, and constructing a reference configuration version for parameter comparison for data storage; acquiring update configuration information by sensing a configuration change event and comparing with a reference version; and comparing the pre-published configuration update information output by the large language model with the actual configuration update information to obtain a parameter change detection result. The invention also relates to a corresponding system, device, processor and storage medium. According to the method, the system and the device for realizing configuration change detection processing based on the large language model, the processor and the storage medium thereof, key configuration information is extracted by starting from a configuration change scheme and utilizing the large model language recognition capability, configuration change content abnormity is accurately recognized, and the configuration change detection processing efficiency is improved. And a production fault risk point caused by configuration change can be effectively detected.
Owner:GUOTAI JUNAN SECURITIES CO LTD

Dynamic sign language time sequence modeling method and system based on attention mechanism

The invention provides a dynamic sign language time sequence modeling method and system based on an attention mechanism, belongs to the technical field of sign language recognition and man-machine interaction, and is used for solving the problems of poor sign language semantic recognition scene adaptability, insufficient personalization and inaccurate semantic feature capture in the related technology. According to the method, hand space coordinates and inertial data are synchronously collected and fused to generate a double-domain time sequence tensor, a four-dimensional fusion feature is constructed in combination with rhythm features and scene prior, semantic expression is enhanced through three-level attention weighting, scene adaptation is achieved based on hierarchical meta-parameters, and the accuracy of scene matching is improved. Personalized optimization is achieved through hand feature fingerprints and full-link parameter feedback, and precise sign language semantic modeling, scene self-adaption and individual adaptation are achieved.
Owner:SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION

Code repair method and device, electronic equipment, storage medium and product

Embodiments of the present application provide a code repair method and device, electronic equipment, storage medium and product. The method comprises: extracting binary features of a packed file, inputting a pre-trained convolutional neural network model to obtain target packing features corresponding to the packed file, performing unpacking processing on the packed file to obtain an unpacked file; performing decompilation processing on the unpacked file to obtain decompiled code, and performing semantic recognition on the decompiled code by a language recognition model to perform symbol recovery processing to obtain symbol recovery code; performing structural processing on the logic of the symbol recovery code by abstract syntax tree technology to obtain structured code; and performing compilation repair processing on the structured code to obtain target code. The above scheme accurately identifies the target packing type by using the pre-trained convolutional neural network model, automatically performs targeted repair according to the target packing type, can accurately repair the influence of the packing tool on the readability of the code, and thus improves the accuracy of code repair.
Owner:SHENZHEN XINGHAN LASER TECH CO LTD

Oil and gas ground construction engineering project auxiliary review method based on natural language processing

The invention relates to the technical field of computer processing, and particularly discloses an oil and gas ground construction engineering project auxiliary review method based on natural language processing, and the method comprises the steps: obtaining a preliminary review material for the preliminary design of an oil and gas ground construction project; performing natural language recognition and content extraction on the pre-audit material based on a pre-trained large model to obtain pre-audit content; performing review analysis on the pre-review content based on a pre-review standard to obtain an analysis result; under the condition that the analysis result is that the pre-examination is passed, obtaining preliminary design data; performing natural language recognition and content extraction on the preliminary design data based on the large model to obtain preliminary design content; and performing review analysis on the preliminary design content based on a design standard to generate a project review result. An existing project review process is optimized by using a natural language recognition technology, so that the review efficiency is improved, and the review comprehensiveness and accuracy are improved.
Owner:SICHUAN SPACE COORDINATE INFORMATION TECH CO LTD

Cross-modal sign language recognition and real-time translation method

The invention discloses a cross-modal sign language recognition and real-time translation method, and the method comprises the steps: fusing the multi-level sign language features of a hand key point, a hand shape, a motion track, a facial expression and a body posture, and combining context semantic understanding and grammatical structure analysis; high-precision continuous sign language recognition and bidirectional translation are realized through sign language grammar structure analysis, a context semantic inference mechanism, continuous sign language segmentation and recognition, a sign language dialect knowledge base and a virtual sign language generator, so that the communication efficiency between hearing impaired people and healthy hearing people is improved. According to the method, isolated sign language vocabularies can be recognized, grammar structures and contexts of sign languages can be understood, bidirectional translation of sign language-text / voice and text / voice-virtual sign language animations is achieved, and recognition and conversion of sign language dialects in different regions are supported.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

Global industry knowledge instant question-answering system based on multilingual pre-training model

The invention relates to the technical field of multi-language natural language processing and knowledge question answering, and discloses a global industry knowledge instant question answering system based on a multi-language pre-training model. Comprising a multi-language input analysis module, a cross-language semantic alignment module, a data preprocessing module, a multi-source knowledge retrieval module, a real-time data crawling and updating module, an answer synthesis module, an answer verification and optimization module, a user interaction module, a historical query storage module and an authority management module. The multilingual input analysis module is used for receiving multilingual questions of a user and carrying out language recognition, word segmentation and preliminary semantic understanding; the cross-language semantic alignment module completes unified mapping of semantics of different languages based on a multi-language pre-training model. Through cooperative work of all the modules, the problems of multilingual semantic deviation, data nonstandard lag and quality safety are solved, multilingual instant and accurate question answering of global industrial knowledge is achieved, and the system is adaptive to multiple scenes.
Owner:BEIJING ZHIYI SHUPU DATA SERVICE CO LTD

Display device and subtitle language identification method

The application provides a display device and a subtitle language recognition method. The method can acquire a subtitle corresponding to a media video in response to a play instruction of the media video, extract visual texture features of the subtitle in a case where the subtitle is an image subtitle, perform language recognition according to the visual texture features to obtain a language recognition result, extract syntax topology features of the subtitle in a case where the subtitle is a text subtitle, perform language recognition according to the syntax topology features to obtain a language recognition result, write the language recognition result into metadata corresponding to the media video, control a display to play the media video and the subtitle, and display a language identifier corresponding to the subtitle according to the language recognition result in the metadata corresponding to the media video. The method recognizes a language by analyzing features of different languages in physical forms and syntax structures, so that the language identifier of the subtitle can be correctly displayed when the media video is played.
Owner:HISENSE ELECTRONICS TECH SHENZHEN CO LTD

A visual multi-modal non-contact gesture unlocking method

The application discloses a visual multi-modal non-contact gesture unlocking method, which comprises the following steps: S1, collecting multi-source sensing data; S2, pre-processing RGB image data; S3, fusing event flow and RGB image, taking the features of fingers and palms as vertices and connecting lines as edges to express the spatial structure of hands; the time correlation is constructed by connecting lines of the same vertices in the front and rear continuous frames as edges, and then the event flow is dynamically mapped into a graph structure to obtain multi-modal fusion features; S4, multi-modal joint prediction, mapping the multi-modal fusion features into a neural network model of a latent space, outputting parameters mu and sigma of a latent distribution, and generating latent features Z through a reparameterization technique; after the latent features Z are superimposed with an additional condition channel, the superimposed features are decoded through a multi-layer perception (MLP) to obtain a final hand posture prediction. The application provides a low-delay, high-robustness real-time gesture interaction and high-precision sign language recognition suitable for challenging environments such as light changes and rapid motions.
Owner:HANGZHOU INST FOR ADVANCED STUDY UCAS

Sign language action recognition and judgment method and system based on multi-channel feature fusion

The invention provides a sign language action recognition and judgment method and system based on multi-channel feature fusion, and the method comprises the steps: extracting hand and body posture features in a video frame, and generating multi-channel feature data; hand and body models of various scenes are loaded, and dynamic switching is carried out to adapt to different recognition requirements; using a machine learning model to predict multi-channel features, and outputting a sign language action recognition result and confidence thereof; and judging the correctness of the sign language action by comparing the recognition result with the expected action. The system solves the problem that an existing sign language recognition system is insufficient in accuracy and robustness under a complex background, and is suitable for various application scenes.
Owner:BEIJING UNION UNIVERSITY

Knowledge graph construction method and device and storage medium

The invention provides a knowledge graph construction method and device and a storage medium, relates to the technical field of knowledge graphs, and can improve the knowledge graph construction accuracy. The method comprises the following steps: acquiring original picture data and original text data; determining an image triple based on the original picture data and a language recognition model; the image triad is used for indicating a plurality of entities in the original picture data and a relationship among the plurality of entities, and the language recognition model is used for determining the image triad; determining a text triple based on the original text data and the large language model; the text triad is used for indicating a plurality of entities in the original text data and a relationship among the plurality of entities, and the large language model is used for determining the text triad; and based on the pre-training model, fusing the image triple and the text triple, and constructing a target knowledge graph.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Visual language recognition method based on spatio-temporal local harmonic neural network and application

The application discloses a visual language recognition method based on a space-time local harmonic neural network and application, and steps of the method comprise the following steps: 1, data preprocessing; 2, constructing a visual language recognition model based on a space-time local harmonic neural network; 3, training of the network model.The method can solve the problems of the existing visual language recognition methods, such as insufficient extraction ability of time characteristics and space characteristics, single method, and not paying attention to the difference of different information, so that the method can accurately recognize the word content in the scene where the speaker posture and the speech speed frequently change, and further provides a new solution for visual language recognition.
Owner:HEFEI UNIV OF TECH +3