Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

389 results about "Language recognition" patented technology

Language Recognition. Summary: The goal of the NIST Language Recognition Evaluation (LRE) series is to establish the baseline of current performance capability for language recognition of conversational telephone speech and to lay the groundwork for further research efforts in the field.

Model for evaluating and predicting mild cognitive impairment risk of old people in nursing institution

The invention relates to a model for evaluating and predicting mild cognitive impairment risk of old people in a pension institution. The model sequentially comprises a behavior analysis module, a language recognition module, a social modeling module, a toughness calculation module, a feature fusion module, a risk reasoning module and the like. Behavior deviation characteristics and abnormal time periods are extracted by collecting behavior data of daily life, diet, social contact and the like of old people and comparing the behavior data with an institution work and rest template; in combination with nursing records, extracting language anomaly features; analyzing social frequency and structure changes in the abnormal time period, and extracting social variation features; a cognitive toughness index is calculated by integrating the health archive and the recovery ability to the health event; and performing toughness weighting on the multi-dimensional features to construct a time sequence tensor, and inputting the time sequence tensor into a recursive model to predict a cognitive impairment risk value. And if the risk value suddenly changes, the system automatically backtracks the feature trajectory of nearly 7 days, constructs and screens a prediction path with the strongest interpretation force, outputs a dominant prediction result and a key factor sequence, and realizes high-interpretability and high-reliability early recognition and intervention reference.
Owner:ZHEJIANG CHINESE MEDICAL UNIVERSITY

Skeleton sign language recognition method of double-flow space-time dynamic graph convolutional network fused with residual learning

The invention discloses a skeleton sign language recognition method of a double-flow space-time dynamic graph convolutional network fused with residual learning, and belongs to the technical field of artificial intelligence and gesture recognition. According to the method, an input gesture skeleton sequence relative to a face is divided into two data streams, namely a hand form data stream and a wrist track data stream through double-reference-system differential homeomorphic mapping; the method comprises the following steps: firstly, processing hand posture data, and capturing a hand joint spatial topological relation by combining a spatial-temporal dynamic graph convolutional network (STDGCNN) with a residual convolutional block; meanwhile, a Finsler trajectory dynamics encoder (FTDE) is adopted to carry out differential geometric modeling on the wrist trajectory, and the direction sensitivity characteristic of the trajectory is captured through a multi-scale causal convolutional network. Then, mutual enhancement of double-flow features is realized through a bidirectional cross feature enhancer (BCFE), and the problem of geometric inconsistency of a heterogeneous feature space is solved through a geometric-driven optimal transmission fusion device (Geometric-OT). The method solves the challenge that a traditional sign language recognition method processes complex space-time correlation of gesture forms and motion tracks at the same time, the technical problems of insufficient relation between hands and faces, insufficient feature expression ability and low space-time feature extraction efficiency, and the problems of geometric inconsistency, single reference system and the like. And the identification accuracy and the real-time performance are obviously improved. Experiments show that the accuracy rate of the method in complex hundreds of sign language vocabulary recognition tasks reaches 95% or above on average, the reasoning speed is only 17ms on average, high-precision real-time sign language recognition is achieved, and the method has higher robustness in complex environments such as noise and shielding.
Owner:刘良锦

Sign language recognition method and device based on multi-modal deep learning

The invention discloses a sign language recognition method and device based on multi-modal deep learning, and the method comprises the steps: multi-modal data input: capturing hand motions, gesture tracks and facial expressions at the same time through a camera, a motion capture sensor and other devices, and forming multi-modal data input; sign language action recognition: precise recognition of sign language actions is realized through a combined model of a deep convolutional neural network and a long-short-term memory network; facial expression and gesture track combined recognition: realizing understanding and translation of complex sign language sentences by combining the captured facial expressions and gesture tracks; and context natural language processing: generating a target statement in combination with context semantic understanding, and outputting a translation result. Through hand motion capture, facial expression analysis, gesture trajectory tracking and context natural language processing, complex sign language motions can be recognized more accurately and translated into characters or voices in real time, and low delay and high accuracy are achieved.
Owner:MIANYANG CITY UNIV

Multimodal brain-computer interface decoding method and related device

The invention belongs to a decoding method, and provides a multi-modal brain-computer interface decoding method and a related device for solving the technical problems that an existing non-intrusive brain language decoding method is insufficient in adaptability in global context and weak in generalization performance in a cross-subject scene, and multi-modal neural feature collaborative enhancement is difficult to achieve. And determining the called execution module. The execution module comprises a feature extraction module, a cross-subject standardization module, a multi-mode semantic fusion module, a language recognition module and a semantic consistency module. By obtaining the unified semantic representation and combining the beam search algorithm, the fairness and universality in different language groups are remarkably improved, multiple modes can be supported, and the brain signal decoding precision and robustness are improved. In addition, cross-subject semantic representation generalization can be realized, and individual specificity is effectively reduced.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Intelligent cockpit implementation method and system based on single-model multi-task reasoning

The invention discloses an intelligent cockpit implementation method and system based on single-model multi-task reasoning, and belongs to the field of automotive electronics, and the method comprises the steps: inputting shared depth features into a detection head and a first classification head after original image data in an intelligent cockpit is preprocessed; the output of the detection head is subjected to detection post-processing to obtain a target detection anchor frame and position information, and the output of the first classification head is subjected to classification post-processing to obtain distraction identification classification; after the target detection anchor frame and the position information are preprocessed, human body area sub-images are input into a safety belt recognition and classification model; inputting the face region sub-images into a fatigue recognition classification model, a face ID matching model, a fixation point estimation model, a sight line estimation model and a head posture estimation model in parallel; and a face image sequence and a voice sequence in the continuous video stream are collected, the face image sequence and the voice sequence are respectively input into the lip language recognition model and the voice recognition model, a recognition result is input into the large language model to be processed to generate an interaction instruction, and a response is output through the voice synthesis module.
Owner:北极雄芯智驾科技(无锡)有限公司

Neck hanging type sign language interpretation equipment and sign language semantic recognition interpretation method thereof

The invention discloses a neck-hung sign language interpretation device and a sign language semantic recognition interpretation method thereof.The neck-hung sign language interpretation device comprises a wearing support which defines a neck surrounding space and comprises a middle connecting arm and two shells, and the two shells are connected to the two sides of the middle connecting arm respectively; the image acquisition unit is arranged on the front side of the shell, comprises a three-dimensional depth camera and an adjusting mechanism for driving the three-dimensional depth camera to rotate, and is used for dynamically tracking and acquiring continuous sign language videos; the processing module is integrated in the shell, comprises a main control processor and a memory, and is used for operating a sign language recognition algorithm; the voice output unit is electrically connected with the main control processor and is used for playing the translation result; and the power supply unit supplies power to the image acquisition unit, the processing module and the voice output unit. Gesture languages used by the deaf-mute for communication are converted into natural languages which can be easily understood by normal people, the deaf-mute group is helped to speak, and the problem of information output when the deaf-mute communicates with the normal people is solved.
Owner:SOUTHEAST UNIV

Information processing method and system for voice conversion

The invention relates to the technical field of voice signal processing and synthesis, and particularly discloses an information processing method and system for voice conversion, and the method comprises the steps: extracting a language feature vector of each word from an input text; calculating posterior probabilities of the words in different languages by combining a Bayesian reasoning mechanism, and generating language attribution confidence coefficient characteristic values; a Monte Carlo sampling method is adopted to carry out multiple times of context sensitive simulation, and pronunciation path selection probability distribution characteristic values are generated; further fusing the feature values into a multi-language pronunciation decision vector, inputting the multi-language pronunciation decision vector into a multi-language end-to-end speech synthesis model, calling a phoneme mapping rule of a corresponding language and an acoustic parameter prediction module, and generating a high-quality target speech spectrogram; and finally, dynamically adjusting Bayesian prior distribution and a language recognition threshold according to the output speech spectrogram.
Owner:SHANDONG POLYTECHNIC COLLEGE

Sign language recognition method based on millimeter wave radar 3D point cloud

The invention relates to a sign language action recognition problem, and provides a sign language recognition method based on a millimeter wave radar 3D point cloud, sign language recognition is carried out in a sequence-to-sequence mode through a non-contact FMCW radar, and the sign language recognition method mainly comprises a first part, namely a continuous sign language action recognition method based on a millimeter wave radar; a second part: establishing an evaluation index STT Score of radar continuous sign language action recognition; according to the invention, a low-cost portable wearable acquisition device is adopted, so that a subject can use the device anytime and anywhere; meanwhile, millimeter-wave radar detection does not depend on light rays, the calculated amount is small, the privacy of a user is effectively protected due to the fact that signals collected by the millimeter-wave radar are radar radio frequency signals, and meanwhile the use feeling is good and the cost is low due to the non-contact collection mode of the millimeter-wave radar.
Owner:ZHEJIANG SCI-TECH UNIV +1

Method and device for generating fine-grained semantic description from action video data

According to the method and device for generating the fine-grained semantic description from the action video data, the training data set is established based on the isolated word sign language recognition data set and the continuous sign language recognition data set containing the word target notation, and the action video data and the action description text data of fine-grained semantic description modeling are obtained; a fine-grained semantic action description style pre-training generation model is obtained through a training framework comprising an action video feature coding module, a multi-modal feature fusion module and a text feature coding module by combining user cue words and system cue words and introducing a mask reconstruction mechanism, action video data is adopted for fine tuning, a loss function is established, and a fine-grained semantic action description style is obtained. The fine-grained semantic action description generation model is used for generating high-quality fine-grained semantic action description data, and the problem that the current fine-grained semantic action description data is insufficient is solved. And the stability and the accuracy of a generated result are ensured when high-dynamic complex scenes such as sign language videos and interactive actions are processed.
Owner:ZHEJIANG UNIV

Sign language recognition glasses device, system and method

The invention discloses a sign language recognition glasses device, system and method, and belongs to the technical field of intelligent equipment, and the sign language recognition glasses device comprises a glasses frame which is used for supporting all components of the device; the camera is arranged on the front side of the glasses frame and is used for capturing hand actions in real time to generate image information; the data processing module is embedded in the glasses frame, electrically connected to the camera and used for processing the image information and executing sign language recognition; the main control unit is arranged on an ear rod of the glasses frame, is electrically connected to the data processing module and is used for coordinating the operation of each component; the display module is arranged at the positions of the lenses of the glasses frame, electrically connected to the main control unit and used for displaying character information obtained after sign language recognition, communication between hearing-impaired people and common people is not limited by specific places any more, and instantaneity and convenience of communication are greatly improved.
Owner:HANGZHOU YIDIAN ELECTRIC TECH CO LTD

Deaf-mute sign language translation pronunciation system

The invention relates to the technical field of sign language intelligent communication, in particular to a deaf-mute sign language translation pronunciation system which comprises a data acquisition and preprocessing module, a sign language model recognition module, a dynamic gesture model module and a voice conversion module. According to the sign language translation pronunciation system for the deaf-mute, wide-angle camera intelligent glasses are adopted for visual collection, wearable equipment is not needed, video streams and thermodynamic diagrams are combined, visual information and key point probability distribution are complementary, the influence of single-mode noise is reduced, video space-time features are extracted through a 3D CNN, long-range dependence is captured in combination with a self-attention mechanism of Transform, and therefore the sign language translation pronunciation system for the deaf-mute is obtained. The sign language recognition precision is improved, an internal reference matrix and a distortion coefficient are calculated through a calibration board, wide-angle lens distortion is corrected, it is ensured that hand key points are accurately positioned, YOLOv8 is adopted to segment an interference object, background noise is prevented from affecting recognition, 21 key point thermodynamic diagrams are generated through MediaPipe, fault tolerance of low-confidence-coefficient key points is enhanced through Gaussian kernel diffusion, and the recognition accuracy is improved. And the model identification precision is improved.
Owner:马也顺

Continuous sign language recognition method based on layered space-time enhancement

The invention discloses a continuous sign language recognition method based on layered space-time enhancement. The method comprises the following steps: acquiring a sign language video; and inputting the sign language video into the trained sign language recognition model to obtain a first recognition result and a second recognition result, and taking the first recognition result as a final sign language recognition result. According to the continuous sign language recognition method based on hierarchical space-time enhancement, multi-stage output of a ResNet34 network is captured through an alignment module, and additional hierarchical alignment supervision is provided; according to the method, the alignment module and the timing causal module are integrated into the ResNet34 network, future and past information is aggregated through the timing causal module, so that more accurate vocabulary boundary perception is achieved, the alignment module and the timing causal module are integrated into the ResNet34 network, good balance between accuracy and calculation cost is achieved, an efficient and accurate solution is provided for a continuous sign language recognition task, and the continuous sign language recognition efficiency is improved. The recognition performance and robustness are remarkably improved, and the problem of under-fitting of the ResNet34 network in space and time contexts is solved.
Owner:ZHEJIANG UNIV OF TECH

Multi-target hand posture estimation method based on millimeter wave radar

The invention belongs to the field of millimeter-wave radar sensing and hand posture estimation, and discloses a millimeter-wave radar-based multi-target hand posture estimation method, which comprises the following steps of: 1, creating a hand region: detecting all objects in a millimeter-wave radar sensing region, detecting a predefined wake-up gesture through a wake-up detection model to recognize a target user, the method comprises the following steps: step 1, evaluating the occurrence probability of a wake-up gesture, and recognizing a hand region, step 2, in the created hand region, constructing a distance-azimuth-pitching three-dimensional feature cube which represents hand posture features; and step 3, estimating hand posture features through an m < 2 > HandNet model. The millimeter wave radar is used as a research object, hand posture estimation is achieved through the millimeter wave radar, multi-user and multi-hand interaction is supported, the shielding problem is effectively solved, and wide application scenes such as intelligent home control, intelligent terminal interaction, virtual reality equipment and sign language recognition are met.
Owner:NANJING UNIV OF POSTS & TELECOMM

Assembly line flow adaptive configuration method and system

The invention relates to the technical field of software engineering, and discloses an assembly line flow adaptive configuration method and system. The method comprises the following steps: obtaining source code data of a software project and performing static analysis and feature extraction to generate structured language description data; querying and matching in a preset template library based on the structured language description data to generate a task template set; performing task dependency relationship analysis and process topology assembly on the task template set to generate initial pipeline data; and performing dynamic strategy adjustment on the execution process of the initial pipeline data by adopting the real-time context data corresponding to the software project to generate target pipeline data. Through automatic language recognition and intelligent process generation, automatic configuration of the continuous integration process is realized, and the technical problem of low automation level caused by lack of automatic recognition and adaptation capability for a project development language environment in a project continuous integration and continuous deployment technology is solved.
Owner:SHANGHAI DONGFANG HOPE SOFTWARE TECH CO LTD

Smart body system for gesture recognition and silent language recognition

An intelligent body system for gesture recognition and silent language recognition comprises a signal acquisition module composed of flexible strain sensors, a signal processing module, a language enhancement module, an augmented reality interaction module and an intelligent auxiliary module. Gestures or lip actions input by a user are recognized, after preliminary classification is conducted through a neural network, a large language model is called to achieve language completion and multi-language translation, an augmented reality interface is used for displaying interaction content, meanwhile, intelligent home equipment can be controlled, and the control state can be fed back in real time. The system is particularly suitable for communication of groups with sound production difficulty, has the advantages of humanization, high recognition rate, strong language complementation ability, natural interaction mode and the like, and realizes a complete intelligent interaction closed loop from perception to understanding and from recognition to control.
Owner:XI AN JIAOTONG UNIV

Deep learning-based sign language-to-multilingual text speech mutual conversion method and application thereof

The invention relates to a sign language-to-multilingual text speech mutual conversion method based on deep learning and application thereof, belongs to the field of artificial intelligence, man-machine interaction and barrier-free communication, and particularly relates to a sign language recognition, semantic understanding and multilingual speech synthesis method based on deep learning and an application system thereof. The sign language to voice text conversion method comprises the following steps: 1, an image acquisition and preprocessing module, 2, a hand key point detection module, 3, time sequence modeling and sign language recognition, 4, natural language understanding and semantic optimization, 5, a multi-language translation module, and 6, a text and voice synthesis output module. According to the sign language-to-multilingual text voice mutual conversion method based on deep learning and the application thereof, sign languages can be automatically, flexibly and accurately converted into voices and characters, and deaf-mutes are helped to efficiently communicate with common people; the system can be widely applied to hospitals, schools, government affair windows, public transportation and other scenes.
Owner:MINXI VOCATIONAL & TECHN COLLEGE

System and method for recognizing early Alzheimer's disease through mobile phone voice

The invention relates to the technical field of medical software, in particular to a system and method capable of recognizing early Alzheimer's disease through mobile phone voice, the system capable of recognizing early Alzheimer's disease through mobile phone voice comprises intelligent mobile phone software and a hospital background, and the mobile phone software builds an interactive platform based on the hospital background. Through intervention monitoring of mobile phone software, when voice communication, chatting and instruction input are carried out in daily life, monitoring can be carried out through system monitoring of software and characteristics such as smoothness and continuity of voice, so that analysis and diagnosis of the Alzheimer's disease can be carried out according to comparison with a normal state, and moreover, in the monitoring process, the monitoring efficiency is improved. Diversified collection is carried out in different time periods and different statement numbers, language recognition and dialect recognition are added when a user inputs voice, and keyword recognition for determining the Alzheimer's disease condition can be carried out through approximate comparison with mandarin after the user speaks out statements.
Owner:GUANGZHOU RED CROSS HOSPITAL

Inference language model training and inference method and model based on microstack and computer equipment

The invention provides a microstack-based reasoning language model training and reasoning method, a model and computer equipment, and the method comprises the steps: calling a to-be-trained model to sequentially carry out multiple times of feature extraction on initial semantic features of a language, and obtaining predicted semantic features; calculating a loss value between the predicted semantic feature and a standard semantic feature of the language; and if the loss value satisfies a preset convergence condition, determining the to-be-trained model as a target inference language model based on a microstack, the target inference language model based on the microstack having a function of identifying semantic features of the language. The to-be-trained model comprises i feature extraction layers, and when the feature extraction layers extract features, data stored in the preset storage unit is added to input data on the basis that only the output data of the previous feature extraction layer exists, so that the accuracy of language recognition is improved.
Owner:PEKING UNIV +1

Multi-language corpus automatic construction and translation optimization system

The invention relates to a natural language processing and multi-language data processing technology, and discloses a multi-language corpus automatic construction and translation optimization system which comprises a corpus collection module, a language recognition and grouping module, a semantic alignment module, a translation optimization module and a corpus quality evaluation and screening module. According to the system, multi-language text data can be automatically collected from the Internet, language recognition and structured storage are carried out, semantic vector coding and alignment matching are carried out on sentences of different languages through a cross-language pre-training model, and high-quality parallel corpus generation is achieved. Meanwhile, the translation model is subjected to incremental training and language proportion regulation and control by utilizing the aligned corpora, so that the translation performance is improved; corpus quality is automatically evaluated and screened through a scoring mechanism, and data reliability is guaranteed. The system can be widely applied to the fields of multi-language machine translation, cross-language information retrieval, intelligent corpus construction and the like.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Automatic test script intelligent generation system

The invention discloses an automatic test script intelligent generation system, belongs to the technical field of software testing, and aims to solve the technical problems of how to realize automatic generation of test scripts, reduce dependence on professional technicians and improve test efficiency. The input module is used for inputting natural languages, extended prompt words, an initial browser and setting parameters; the intelligent analysis layer is used for realizing multi-language environment automatic test execution by combining the multi-language recognition capability of a large model and manual test cases of a set of languages; the element recognition layer is used for supporting a CSS / visual / semantic hybrid positioning strategy in combination with the visual recognition interface element capability and automatically selecting an optimal positioning scheme; and the execution feedback layer is used for acquiring latest page element information in real time when the UI structure is changed, and automatically utilizing the page semantic comprehension capability of the large model and the visual recognition capability of the multi-modal large model.
Owner:INSPUR QILU SOFTWARE IND

Visual multi-mode non-contact gesture unlocking method

The invention discloses a visual multi-mode non-contact gesture unlocking method. The method comprises the following steps: S1, collecting multi-source sensing data; s2, preprocessing the RGB image data; s3, the event stream and the RGB image are fused, the features of the fingers and the palm are used as vertexes, connecting lines are used as edges, and a hand space structure is expressed; constructing time correlation by taking a connecting line of the same vertexes in front and back continuous frames as an edge, and dynamically mapping an event stream into a graph structure to obtain a multi-modal fusion feature; s4, performing multi-modal joint prediction, mapping the multi-modal fusion feature to a neural network model of a potential space, outputting parameters mu and sigma of potential distribution, and generating a potential feature Z through a re-parameterization technique; and after the potential feature Z and the additional condition channel are overlapped, decoding is carried out through a multi-layer perceptron (MLP), and final hand posture prediction is obtained. The invention provides low-delay and high-robustness real-time gesture interaction and high-precision sign language recognition suitable for challenging environments such as illumination variation and fast movement.
Owner:HANGZHOU INST FOR ADVANCED STUDY UCAS

Sign language-voice conversion system

The invention discloses a sign language-voice conversion system, and belongs to the technical field of auxiliary communication and wearable computing. The system comprises a wearable myoelectricity acquisition module used for acquiring double-arm myoelectricity signals when a user executes sign language; the mobile terminal module is wirelessly connected with the acquisition module and is used for receiving and preprocessing the signal and uploading the signal; the cloud processing module is used for receiving the signal, converting the signal into text information through a sign language recognition model, and further calling a voice synthesis service to convert the text into voice data; and the wearable audio output module is used for receiving and playing the voice data. Through an innovative end-to-end hardware system architecture, natural, accurate and real-time translation and voice output of sign language gestures are realized, communication barriers between hearing-impaired people and healthy hearing people are effectively solved, and the system has the advantages of flexible deployment, user friendliness and privacy protection.
Owner:宋飞 +1

Chinese lip language recognition method based on credible visual position element acquisition

The invention discloses a Chinese lip language recognition method based on credible visual position element acquisition. The method comprises the following steps: S1, data acquisition and preprocessing: obtaining video data depicting lip motion; s2, deep clustering: performing deep clustering on the lip motion depicting video data to obtain the number of cluster-distributed visual position classes, corresponding visual position classes and a visual position library, so as to obtain frame-by-frame image data with visual position class labels corresponding to the lip motion depicting video data; and S3, identifying the cascade Chinese character sequence based on the visual position element intermediate representation: performing feature extraction based on the frame-by-frame image data with the visual position element category label to realize the identification of the cascade Chinese character sequence taking the visual position element as the intermediate representation. According to the method, the cumulative error of recognition prediction can be reduced, the lip language recognition performance based on the visual position elements is improved, and the accuracy bottleneck of the lip language recognition based on the visual position elements is broken through.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Multi-mode anti-interference communication method and system based on lip language recognition

The invention discloses a multi-mode anti-interference communication method and system based on lip language recognition, and belongs to the technical field of communication equipment. The method comprises the following steps: acquiring a face lip video stream and an audio signal; in response to the conventional mode trigger signal, feature extraction is performed on the lip video stream and the audio signal, and extraction results are fused to generate a fused feature vector; performing voice enhancement on the fusion feature vector in combination with lip motion information, and outputting an audio enhancement signal; and in response to the silent communication mode trigger signal, performing lip language recognition based on the face lip video stream to obtain a lip language recognition text, and converting the lip language recognition text into voice. Clear and stable communication in an ultra-strong noise environment can be realized by combining two types of modal information, and the problem that the communication quality is influenced by an existing high-noise environment and the requirement for mobile silent communication in a special scene are solved.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Full-language voice interaction legal affair agent system and method for small and micro enterprises

The invention discloses a full-language voice interaction legal affair agent system and method for small and micro enterprises. The core of the system comprises a full-language voice input module, a language recognition module, a voice recognition engine routing module, a special voice recognition engine cluster, an enterprise-level multi-language natural language processing module, an enterprise legal affair agent core, a multi-mode output module and a feedback optimization module, and all the modules work cooperatively. And seamless connection from voice input to scheme output is realized. The process of the method comprises voice acquisition and language recognition, audio routing and text conversion, semantic understanding and intention recognition, commercial risk multi-dimensional trade-off analysis, structured scheme generation and multi-modal output, and meanwhile, the system is driven to continuously iterate through a feedback optimization mechanism. According to the method, the working efficiency and experience of an enterprise owner are remarkably improved, the legal consultation revolution of'moving without using hands' is achieved, the method has low threshold and high specialty, legal instruments can be directly generated, and the method has high intelligence and self-adaptive capacity.
Owner:齐洪建

Method and system for supporting multilingual environment of voice instruction

The invention relates to the technical field of voice instruction recognition, and provides a multilingual environment support method for a voice instruction, which comprises the following steps: S1, sending a command word acquisition request to a firmware end by an applet end, and acquiring multilingual basic command words and wake-up word configurations stored in the firmware end; s2, selecting a to-be-modified target command word, collecting multilingual voice data through a microphone at a small program end, and uploading the voice data to a cloud end; s3, the cloud calls a multi-language multi-command recognition model to train and deduce the voice data; s4, the cloud end transmits the command word configuration data back to the applet end, and the applet end transmits the multilingual command word configuration data to the firmware end through Bluetooth; and S5, the firmware end receives the command word configuration data and updates the command word, so that the device supports multilingual recognition and response to the new command word. A multi-language model is dynamically trained through a cloud end, and configuration is issued, so that multi-language seamless switching of the equipment and flexible adjustment of wake-up words and command words are realized.
Owner:SHANGHAI SHENSILICON SEMICON CO LTD

Realtime AI sign language recognition with avatar

Disclosed herein are method and system aspects for translating between a sign language and a target language and presenting such translations. For example, a method receives input language data and translates the input language data into sign language grammar. The method retrieves phonetic representations that correspond to the sign language grammar from a sign language database and generates coordinates from the phonetic representations using a generative network. The phonetic representations are digital representations of individual signs created through manual input of body configuration information corresponding to the individual signs. Further, the method renders an avatar that moves between the coordinates. In another example, a bidirectional communication system allows for realtime communication between a signing entity and a non-signing entity.
Owner:SIGN-SPEAK INC

Sound environment analysis and monitoring method based on artificial intelligence

The invention discloses a sound environment analysis and monitoring method based on artificial intelligence, and belongs to the field of artificial intelligence. Sound feature extraction and parameter analysis; converting sound into text; converting the sound event into a text description form by using a multi-mode big language model with a sound event analysis function and outputting text information; timestamps are added to the output text information, and the text information is classified, sorted, recorded and stored for a user to trace back; determining a key sound event by identifying a dialogue scene; after the key sound event is triggered, the device reminds the user according to a specific form. Any equipment does not need to be implanted, and postoperative risks and maintenance cost do not exist; a user can perceive environment sound and understand dialogue content without listening by himself or herself, and a hearing aid is not needed; sign language actions do not need to be captured through a camera, and the problem of dialect sign language recognition errors does not exist; the user does not need to stare at the character display device in real time during use, and the user is reminded in a specific mode after the key sound event is triggered.
Owner:SHENZHEN TECH UNIV

Human body action and language instruction combined recognition system

The invention discloses a human body action and language instruction combined recognition system, and particularly relates to the field of intelligent recognition, which comprises a language recognition module, an action recognition module, a mutual information value module, an independent analysis module, a fusion analysis module and an instruction generation module, the language recognition module is used for collecting human body language information and performing feature extraction on the human body language information to generate a language signal X; the action recognition module is used for collecting human body action information and performing feature extraction on the human body action information to generate an action signal Y; the mutual information value module constructs joint distribution from the language signal X and the action signal Y, and obtains probability distribution of the language signal X, probability distribution of the action signal Y and joint probability distribution of the language signal X and the action signal Y. According to the system, stable recognition can still be achieved in the complex environment with various accents and various visual angles, interaction errors caused by signal misjudgment are reduced, and a user can obtain natural and efficient experience in different interaction scenes.
Owner:PUWANG (SHANGHAI) INFORMATION TECH CO LTD

Tablet computer real-time voice recognition and translation system based on side cloud collaboration

The invention discloses a tablet computer real-time speech recognition and translation system based on side cloud collaboration, which relates to the technical field of speech processing, and comprises a speech noise reduction module for performing multi-stage enhancement by adopting beam forming and a deep residual network, performing multi-channel feature fusion in combination with adaptive noise estimation and an attention mechanism, and performing speech recognition and translation. Obtaining clean voice data after signal-to-noise ratio optimization; the language recognition module inputs the clean voice data into a language recognition network, extracts a voice feature vector, performs language recognition through a language clustering model and generates a language tag; and the translation module is used for carrying out semantic optimization through semantic understanding and context modeling according to the preliminary recognition text, and carrying out translation processing by utilizing a neural network translation model to generate a translated text. According to the invention, the technical effect of effectively improving the signal-to-noise ratio of the voice signal in a complex acoustic environment is realized.
Owner:GUANGDONG OUDULIFANG TECH CO LTD