Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

358 results about "Confidence measures" patented technology

Confidence-building measures. Confidence-building measures (CBMs) or confidence- and security-building measures are actions taken to reduce fear of attack by both (or more) parties in a situation of tension with or without physical conflict.

Motor fault detection method and system based on voiceprint recognition

The invention discloses a motor fault detection method and system based on voiceprint recognition. According to the method, an annular microphone array is adopted to collect motor sound signals in a non-contact mode, a three-channel time-frequency data set is constructed through empirical mode decomposition (EMD) and a Mel-frequency cepstral coefficient (MFCC), fault diagnosis is carried out in combination with a CNN + ResNet network, and dynamic time warping (DTW) and CNN fusion matching is supported. The system comprises a preprocessing module, a fault template library and a matching algorithm, integrates wavelet denoising and multi-beam acquisition technologies, covers a frequency band of 50Hz-20kHz, can display a fault type and trend analysis in real time, and triggers secondary verification when the confidence coefficient is insufficient. According to the scheme, the anti-interference capability is improved through array signal processing, model parameters are optimized in combination with transfer learning, non-contact detection is achieved, the real-time performance and accuracy of fault diagnosis are remarkably improved, and the method is suitable for industrial motor health monitoring.
Owner:GUANGZHOU DAYIN ZHIYUAN DIGITAL TECH CO LTD

Cement equipment maintenance decision-making method and device based on knowledge graph and large model reasoning

The invention provides a cement equipment maintenance decision-making method and device based on a knowledge graph and large model reasoning, relates to the field of cement industry intelligent operation and maintenance, and solves the technical problem of decision-making response delay caused by knowledge fragmentation. The method comprises the following steps: extracting real-time characteristics from vibration spectrum signals, temperature curves and torque waveform data collected by an edge gateway, and extracting a work order entity triple from a natural language work order text of an EAM system; based on an equipment BOM list, a historical maintenance record and an FMEA analysis table, physical assembly constraint conditions are defined through ontology modeling to generate a cement equipment topological relation and a fault rule chain, and a knowledge graph is created to output a fault rule base with confidence coefficient weights. And inputting the real-time feature vector and the work order entity triple into a multi-modal collaborative inference engine, triggering a matched fault rule chain by combining real-time features and semantic features, outputting a fault root cause and an associated maintenance strategy ID, and labeling a logic chain. And activating the associated maintenance strategy ID and obtaining the real-time characteristic deviation degree of the maintenance strategy ID, quantifying the decision credibility through a tracing rule matching path, obtaining an executable maintenance instruction packet with a logic chain, executing the maintenance instruction packet and dynamically updating the knowledge graph based on a maintenance result. The method is used in the maintenance decision-making process of the cement equipment.
Owner:HEFEI CEMENT RESEARCH AND DESIGN INSTITUTE CO LTD

Knowledge base question and answer platform construction method based on large language model

The invention relates to the technical field of natural language processing, and discloses a knowledge base question and answer platform construction method based on a large language model, which comprises a knowledge acquisition module, a data preprocessing module, a text processing module, a vectorization module, a question understanding module, a mixed retrieval module, a prompt generation module, an answer generation module and an answer quality analysis module. A secondary inquiry processing module and a feedback learning module; according to the method, a semantic segmentation algorithm is combined with semantic retrieval and keyword retrieval, so that the flexibility is high; normalized prompts are constructed, input is performed according to correlation sorting, and the accuracy of answers is improved; multi-dimensional confidence evaluation is introduced, strict multi-layer security and compliance filtering is set, and the reliability of the system is ensured; the relevance of multiple rounds of dialogues is judged and complemented, so that interaction is more natural and efficient; knowledge is collected and updated in real time, a knowledge base and a retrieval strategy are continuously optimized, and a closed loop of data-application-feedback-tracing-optimization is formed.
Owner:JIANGSU INSPIRE INTERNET OF THINGS TECH CO LTD +1

Transient wave recording fault intelligent judgment method combining waveform matching and cloud edge cooperation

The invention provides a transient recording fault intelligent judgment method combining waveform matching and cloud edge cooperation, and relates to the technical field of fault diagnosis. Comprising the following steps: collecting global topological information and fault judgment reference parameters of a master station synchronous power system, and outputting initial transient recording data; carrying out preprocessing and feature extraction on the initial transient recording data, realizing edge side layered fault detection by adopting an improved waveform matching algorithm, and outputting multi-layer fault information; constructing a global fault judgment model, updating the global fault judgment model through the hierarchical model, and optimizing the fault judgment model; multi-level fault information is input into the optimized fault judgment model for accurate fault judgment, and graded response measures are taken for real-time judgment results with different confidence degrees; according to the invention, through deep fusion and system-level optimization of multiple technical streams, a transient recording fault intelligent judgment system integrating real-time performance and accuracy is created.
Owner:ZHUHAI GANXING AUTOMATION EQUIP CO LTD

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Small-sample contrast enhancement fine tuning method and system based on large language model

The invention relates to a small-sample contrast enhancement fine tuning method and system based on a large language model, which are used for identifying named entities in recruitment texts. The method comprises the steps of performing cleaning and format conversion on an original recruitment text, and generating an input sample conforming to a natural language instruction format; under the condition that the labeled samples are insufficient, positive and negative sample pairs are constructed to enhance the recognition capability of the model on entity categories and boundaries; carrying out low-rank parameter updating on the pre-trained large language model by adopting a LoRA fine tuning technology, and reducing computing resource consumption in combination with 4-bit quantitative training; in a pre-training large language model reasoning process, through a multi-dimensional joint confidence evaluation mechanism, confidence of four dimensions of entity levels, lengths, types and contexts is synthesized, and low-confidence identification results are filtered after dynamic weighted normalization processing. The method is suitable for recruitment recommendation, talent matching and other downstream tasks, and has the advantages of high recognition accuracy, low training cost, high system robustness and the like.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Voice interaction processing method, device and system, intelligent door lock and cloud server

The invention is suitable for the technical field of intelligent door locks, and provides a voice interaction processing method, device and system, an intelligent door lock and a cloud server, and the method comprises the steps: receiving a voice instruction sent by a user, and carrying out the local recognition, so as to obtain a local recognition result and the confidence of the local recognition result; when the local recognition result is matched with the local instruction set and the confidence is high, executing a local operation corresponding to the local recognition result; otherwise, establishing a secure communication session with the cloud server based on voiceprint biological characteristics extracted from the voice instruction; sending the voice instruction to a cloud server through the secure communication session, and receiving a cloud operation instruction and / or a dynamic response text returned by the cloud server; and executing corresponding operation according to the cloud operation instruction, and / or converting the dynamic response text into voice for broadcasting through a text-to-voice engine. The intelligent door lock voice interaction system solves the problem that the real-time performance, the reliability, the safety and the interaction intellectualization are difficult to balance in the existing intelligent door lock voice interaction technology.
Owner:SHANGHAI ZHENGZHI INTELLIGENT TECH CO LTD

Multi-modal emotion recognition method and system for service-oriented robot

The invention belongs to the technical field of artificial intelligence, and particularly relates to a service-oriented robot-oriented multi-modal emotion recognition method and system, and the method comprises the steps: collecting audio and video stream data of emotion changes of a user, and separating visual and voice data; extracting visual and voice emotion features through a pre-training model, and calculating prediction probability distribution of each mode; constructing a bimodal confidence quantitative model based on the distribution to obtain each modal confidence; and fusing the features by adopting a sectional type dynamic weight distribution strategy so as to identify the emotional state of the user. Visual and voice modes are fused, feature alignment is realized in combination with dynamic time warping, spatial optimization performance is shared and expressed through a confidence model, a dynamic weight strategy and a cross-modal time sequence cooperation module, and the method has high recognition accuracy, high robustness and real-time processing capacity in a complex environment and is suitable for various service scenes.
Owner:SUZHOU CITY UNIV

Conference window conversation system based on voice sensor

The invention discloses a meeting window call system based on a voice sensor, and relates to the technical field of communication. Comprising a full-space voice acquisition module, a dynamic confidence evaluation module, a phase coupling enhancement module, a multi-dimensional speech mode recognition module, a dynamic gain adjustment and tone quality enhancement module and a full-link voice integrity verification module, an array type voice sensing unit is deployed based on audio reflection characteristics, a multi-band full-space acquisition matrix is constructed, voice signals of two parties meeting are acquired, and a basic audio data stream with a high signal-to-noise ratio is output. According to the invention, through array sensing acquisition, dynamic confidence evaluation, phase coupling and harmonic enhancement, multi-dimensional speech recognition and dynamic gain adjustment, accurate recognition and fidelity enhancement of low-volume speech are realized, and through combination of full-link integrity verification and mistaken killing backtracking, weak signals are ensured not to be missed and distorted, so that the accuracy of speech recognition is improved. And the privacy, the security and the information integrity of the meeting call are improved.
Owner:HANGZHOU HUA TING TECH CO LTD

Remote interview anti-fraud system and method based on multi-mode depth forgery detection

The invention discloses a remote interview anti-fraud system and method based on multi-modal deep counterfeiting detection, and relates to the technical field of data analysis. According to the method, face hash and voiceprint hash are generated by collecting identity document information and biological characteristics of job seekers, and after zero-knowledge identity matching is achieved, hash triples are stored in a block chain; in the interview process, three layers of confidence coefficients are calculated by analyzing physiological feature pairs, deep forging abnormal indexes, voiceprint synthesis feature values and lip sound maximum offset, and weights are adjusted based on historical attack rates to generate fusion scores; optimizing an RNN behavior analysis network, outputting a real-time abnormal behavior index, obtaining a final risk score in combination with the fusion score, and triggering a three-level response mechanism; according to the method, the fraud interception rate is increased, the misjudgment rate is reduced, and the problems of identity false use, deep counterfeiting and behavior fraud are effectively solved on the premise of ensuring privacy security.
Owner:UNIV OF SCI & TECH OF CHINA

Off-line and on-line mixed use method of AI voice

The invention relates to the technical field of AI voice, and particularly discloses an off-line and on-line mixed use method of AI voice, which comprises the following six steps of: evaluating a network connection state to obtain a network reliability index; analyzing voice data input by the user and preamble information, determining the complexity of voice intention and generating a demand priority; and then cooperatively selecting a voice processing mode according to the network reliability index, the demand priority and the decision rule. If the mode is an off-line mode, extracting a voice feature vector, classifying responses by using a local model, obtaining a processing mode confidence coefficient according to user feedback, and if the processing mode confidence coefficient is lower than a threshold value, triggering an on-line supplementary response and updating a rule; and if the mode is an online mode, feature vectors are extracted to be matched with a cloud knowledge base, a fine classification result is obtained, and accurate response is performed by means of a cloud large model. The whole process is combined with network conditions and user requirements, and a processing mode is flexibly decided, so that high-quality AI voice interaction service is provided.
Owner:SHENZHEN MAICHIRUI SOFTWARE CO LTD

Intelligent voice question-answering system and knowledge reasoning method for operation and maintenance of power equipment

The invention discloses an intelligent voice question-answering system and a knowledge reasoning method for operation and maintenance of power equipment. The system comprises a knowledge base construction module and a core question and answer engine module, the knowledge base construction module is used for constructing a power equipment knowledge graph and a fault case library, and the core question and answer engine module comprises a natural language understanding unit, a multistage semantic retrieval unit and a voice interaction ambiguity resolution unit. A natural language understanding unit identifies intent categories and power entities. The multi-stage semantic retrieval unit executes the following steps: matching in an ontology layer of the knowledge graph; performing extended query to obtain detailed information; and similar case matching is carried out in the fault case library. The voice interaction ambiguity resolution unit is used for executing disambiguation processing; analyzing the confidence and clarifying the dialogue; and triggering knowledge retrieval and reasoning and generating a sorting answer. The operation and maintenance efficiency is improved, the dependence on experts is reduced, the operation and maintenance quality and safety are improved, knowledge precipitation and sharing are promoted, the field work experience is improved, and scientific decision making is assisted.
Owner:安徽明生恒卓科技有限公司

Intelligent education robot question answering system based on voice recognition and knowledge graph

The invention discloses an intelligent education robot question answering system based on voice recognition and a knowledge graph, and particularly relates to the technical field of artificial intelligence. The method comprises the following steps: performing feature extraction on a user voice signal to obtain a text sequence and a multi-modal context feature; recognizing subject domain judgment and question answering intentions based on supervised classification and keyword rule fusion, and outputting a subject domain prior probability and a question answering target vector; determining polysemy words in the text sequence by using a context window, generating a semantic item sequence and semantic item confidence, constructing a subject domain sub-graph based on a subject domain prior probability, and obtaining a candidate reasoning path and a path scoring vector; determining a target teaching concept and an optimal reasoning path by combining Bayesian inference and consistency verification; generating a personalized question answering result aiming at the question asked by the user through fact retrieval and knowledge derivation in combination with the question answering target vector; accurate and efficient intelligent teaching question answering can be realized, and question answering accuracy and intelligent interaction capability of the education robot are effectively improved.
Owner:SHANDONG BAIKU EDUCATION TECH CO LTD

Text processing method and system based on large language model

The invention discloses a text processing method and system based on a large language model, and relates to the technical field of computers, and the method comprises the steps: collecting text data, generating a variant sample through a data enhancement technology, obtaining a pre-training data set, introducing a knowledge graph, and outputting a trained large language model; performing hyper-parameter optimization by using an adaptive task selector, an incremental learning framework and a meta-learning algorithm; non-text data is considered, different types of context clues are captured by using a multi-modal model, and an expression mode is understood by using a cross-culture cognition framework; based on the updated large language model, establishing a confidence evaluation method, and outputting a high-confidence text processing result; according to a high-confidence-coefficient text processing result, a personalized user portrait is constructed, and recommended content is optimized; according to the method, the multi-modal model and the cross-culture cognition framework are introduced, so that the performance of the large language model in processing complex real world problems is effectively improved.
Owner:BEIJING TAIHE GUANFU TECH CO LTD

Internet enterprise multi-mode identity verification method and system

The invention discloses an Internet enterprise multi-mode identity verification method and system, and belongs to the technical field of Internet enterprise security, and the method comprises the steps: obtaining user historical behavior data, current transaction request data, equipment environment parameters and initial biological signal data, carrying out the risk assessment, and generating a verification path; obtaining a personalized verification sequence instruction, collecting a user face dynamic video stream, a real-time voice stream and response action time sequence data, carrying out cross-modal association comparison with a user reference biological feature template, outputting a biological feature confidence matrix, carrying out association analysis in combination with the obtained structured identity feature vector and risk assessment, and obtaining a personalized verification result; and generating a verification decision feature vector to judge a verification result state, and obtaining a pass instruction, a rejection instruction or a manual auditing request instruction. According to the method, dynamic risk-driven multi-modal verification path generation, cross-modal biological feature association decision and incremental learning mechanisms are adopted, so that the optimal balance between security and user experience can be realized in a complex network environment.
Owner:NAN JING OU YI TAI XIN XI KE JI YOU XIAN GONG SI

Methods and system for providing a data element from a corpus of data

A method and corresponding systems for providing a mapping of an unstructured corpus of data onto a predetermined data structure are provided. The method comprises obtaining a prompt for providing the data element; accessing a corpus of data; providing a machine-learned function configured to identify data elements in corpora of data based on prompts; applying the machine-learned function to the corpus of data to identify at least one data element in the corpus of data corresponding to the prompt; determining a confidence measure for the identified at least one data element using a verification function, the verification function being independent from the machine-learned function; and providing the identified at least one data element as the data element based on the confidence measure.
Owner:SIEMENS HEALTHINEERS AG

Operation interaction method and system applied to camera image editing

The invention discloses an operation interaction method and system applied to camera image editing, and relates to the technical field of image processing, and the method comprises the steps: obtaining a voice semantic heat map, a pointing intensity map, a touch confidence map and a gazing confidence map based on a multi-modal interaction data packet, and calculating an image feature matrix at the same time; fusing into a multi-modal evidence graph through a normalized scale; performing semantic segmentation according to the image feature matrix to obtain a semantic segmentation first draft and a pixel-by-pixel category confidence coefficient, and performing position correlation weighting on the pixel-by-pixel category confidence coefficient by taking the multi-modal evidence graph as a confidence coefficient modulation factor to generate a candidate object mask sequence; and performing highlight display on the candidate object mask sequence, and performing conflict resolution and priority rearrangement in combination with the multi-mode evidence graph to generate a target object mask. According to the method, deep fusion of the interaction intention and image segmentation is realized, the precision and consistency of candidate region detection are improved, and the stability of real-time rendering and the reliability of an editing result are improved.
Owner:SHENZHEN XUJING DIGITAL TECH CO LTD

Electric power operation ticket intelligent generation system based on OCR and voice recognition

The invention discloses an intelligent power operation ticket generation system based on OCR and voice recognition, and relates to the technical field of intelligent generation of power operation tickets, and the system obtains an operation task image, a voice instruction, equipment data and historical records through an OCR terminal, a voice input collector, an equipment monitoring sensor and a rear-end interface; carrying out vectorization and semantic alignment on the multi-source data, and establishing a standardized data structure; establishing a strategy guiding agent model, analyzing the semantic and path reasonability of the operation order, and carrying out intelligent evaluation and decision making; semantic consistency is detected in real time, an equipment connection structure is monitored, and path complexity is calculated; the strategy confidence coefficient is evaluated through multi-chain mapping and consensus analysis; and finally, generating an operation order conforming to the standard and outputting the operation order in an electronic format, thereby ensuring the standardization and traceability of the operation process.
Owner:GUANXI POWER GRID CORP HEZHOU POWER SUPPLY BUREAU

Systems and processes for multifactor authentication and identification

A system can include a server having a processor configured to receive, from a first computing device, first data associated with at least one historical data attack. The processor can encode the first data into at least one fixed-size representation. The processor can generate at least one anonymized vector representation based on the at least one fixed-size representation. The processor can receive, from a second computing device, second data associated with a potential attacker. The processor can encode the second data into at least one additional fixed-size representation. The processor can generate a second anonymized vector representation based on the at least one additional fixed-size representation. The processor can generate a confidence measure of the potential attacker as an attacker based on a comparison between the at least one anonymized vector representation and the second anonymized vector representation.
Owner:T STAMP INC

Voice transfer text error correction method and device, storage medium and computer equipment

PendingCN121706770ASemantic analysisSpeech recognitionAlgorithmTransliteration
According to the voice transliteration text error correction method and device, the storage medium and the computer equipment, after the voice transliteration text is received in real time, the text is subjected to primary error detection, and the first error confidence coefficient is obtained; when the first error confidence coefficient is smaller than a fast error correction threshold value, the voice transcription text is directly output, and unnecessary error correction is avoided; otherwise, performing word-level rapid error correction on the voice transcription text to obtain a rapid error correction text, and performing secondary error detection on the rapid error correction text to obtain a second error confidence coefficient. Judging whether the second error confidence is smaller than a depth error correction threshold value or not; if yes, the rapid error correction text is directly output, and if not, semantic-level deep error correction is conducted on the rapid error correction text through a deep error correction model, and a deep error correction text is obtained and then output. Through a hierarchical error correction processing strategy, the method can be adapted to different service scenes, and meanwhile, dynamic balance between quality and efficiency of different service scenes can be realized by adjusting two error correction thresholds.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

Real-time voice stream dialogue interaction method and system based on large language model

The invention discloses a real-time voice stream dialogue interaction method and system based on a large language model, and belongs to the technical field of voice recognition, natural language processing and human-computer interaction. The method comprises the following steps: S1, segmenting a voice signal according to a 320 millisecond time window; s2, through double-channel parallel processing, generating a preliminary text hypothesis locally and uploading a frame to a server side to accurately reconstruct a text; s3, fusing to generate a final text with a confidence label; s4, inputting the text into a large language model with a context management mechanism to generate a natural language response; s5, performing speech synthesis output of rhythm perception; s6, updating a dialogue cache and synchronizing a language model state; s7, predicting the semantic trend in advance by adopting a delayed slow release mechanism; and S8, introducing round-level intonation / rhythm parameters for semantic adjustment. The method has the beneficial effects that the response delay is reduced, the recognition accuracy is improved by fusing the confidence label, and meanwhile, the continuity and naturalness of multi-round interaction are enhanced through context and intonation modeling.
Owner:MINIMALIST INTERNET (BEIJING) INFORMATION TECHNOLOGY CO LTD

Television terminal education application multi-mode user authentication method and system

The invention relates to the technical field of data processing, and discloses a television terminal education application multi-mode user authentication method and system. The method comprises the following steps: monitoring a user face orientation offset degree and a voice interaction frequency through an IPTV set top box to obtain a learning participation degree index, comparing the index with a preset threshold 0.7 to execute a concentrated learning or distraction state authentication process, and dynamically adjusting through an authentication severity regulator to obtain a high-sensitivity or low-sensitivity authentication parameter, and performing weighted fusion calculation on the multi-modal biological characteristics of the user to obtain an identity confidence score associated with a learning state, and performing matching verification on the identity confidence score and a dynamic authentication threshold to obtain a user identity authentication result based on the learning concentration degree. According to the method and the device, the problem that a user authentication strategy in television terminal education application cannot adapt to dynamic change of a learning state is solved, and the accuracy of multi-modal user identity authentication based on learning concentration and the adaptability of user learning experience are improved.
Owner:CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD +1

Fan blade running state monitoring system and method based on voiceprint recognition

The invention discloses a system and a method for monitoring the running state of a fan blade based on voiceprint recognition. The method comprises the following steps: acquiring voiceprint data and environment data during running of the fan blade through a data acquisition module; performing noise reduction processing by adopting an improved VMD method, and adaptively optimizing a decomposition modal number and a penalty factor through fuzzy entropy and an energy ratio; constructing a plurality of branch input data according to the voiceprint data; inputting the two types of branch data into a pre-trained multi-branch model, carrying out weighted fusion on similar results by combining a voiceprint signal-to-noise ratio and a defect confidence difference, and outputting a state result of a defect type and confidence; and when the confidence exceeds a threshold value, integrating the environment data, the blade parameters and the GAP features to generate positioning analysis data, determining a defect position through a defect positioning model, and triggering an alarm. According to the method, the feature quality is improved through adaptive noise reduction, the feature fusion effect is optimized through dynamic weight, the defect recognition accuracy is enhanced through confidence fusion, efficient positioning is achieved, and the monitoring efficiency and reliability are remarkably improved.
Owner:JIANGSU FRONTIER ELECTRIC TECH

Security risk intelligent response method and system

The invention discloses a security risk intelligent response method and system. The response method comprises the following steps: acquiring video streams, equipment states and environment sensor data in real time through a multi-source interface; processing through a feature fusion engine to generate a threat matrix containing threat type judgment and confidence; matching an emergency plan library based on a dynamic weighting algorithm; an equipment control instruction queue is generated according to a matching result, conflict detection is executed, and multiple instructions of the same equipment are sequenced according to fire control, access control locking and video tracking; and finally, acquiring an equipment execution state in real time and feeding back the state to a plan library to form closed-loop control, thereby solving the problems of high false alarm rate and response lag in traditional security and protection. Historical and real-time data are fused through a dynamic decision model, the limitation of a static plan is broken through, and the disposal efficiency is improved; instruction conflicts are eliminated through a conflict resolution mechanism; and a self-healing system and model correction are combined to form double-closed-loop optimization.
Owner:ZUNYI HUIFENG INTELLIGENT SYST

Intelligent control method and device, electronic equipment and readable storage medium

PendingCN121956605AImprove Fusion Accuracyimprove accuracyComputer controlModal dataEngineering
The invention discloses an intelligent control method and device, electronic equipment and a readable storage medium, and belongs to the technical field of intelligent control. Under the condition that the gesture instruction does not conflict with the voice instruction, performing multi-dimensional normalization on the gesture feature vector to obtain a to-be-fused gesture vector; performing normalization processing on the voice feature vector based on context semantic perception to obtain a to-be-fused voice vector; fusing the to-be-fused gesture vector and the to-be-fused voice vector to obtain a fusion vector, and determining a fusion instruction corresponding to the fusion vector as a target execution instruction; and under the condition that the gesture instruction conflicts with the voice instruction, selecting a target execution instruction according to a first confidence coefficient corresponding to the gesture instruction and a second confidence coefficient corresponding to the voice instruction. The self-adaptive normalization strategy is designed for different modal data, the data fusion accuracy is improved, meanwhile, the gesture and voice instruction conflict can be intelligently and flexibly processed, and the accuracy of determining the real intention of the user is improved.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1

Education robot voice signal processing method

The invention discloses a voice signal processing method for an education robot, relates to the technical field of voice signal processing, and aims to solve the problems of multi-person overlapping language and environmental noise interference in a classroom, multi-channel data is acquired by relying on a microphone array and a camera, and an environmental model is constructed in the step 1 to determine a noise baseline and student distribution; in the second step, overlapped voices are detected, and sound source localization is carried out in combination with the time difference of arrival and mouth shape data; in the third step, directional gain is executed in the target direction, and a deep network is used for separating aliasing voice; and in the step 4, the separated voice is input into a children customized recognition engine to complete high-precision recognition and interaction in combination with confidence evaluation. The recognition accuracy and the interaction efficiency can be remarkably improved under the complex scenes of classroom reverberation and simultaneous speaking of multiple persons, meanwhile, the noise change is tracked through the global environment model so that the education robot can keep stable recognition performance in diversified teaching interaction, and the teaching effect is remarkably enhanced.
Owner:北京爱宾果科技有限公司

Audio identification method and device based on marking and backtracking correction

The invention relates to an audio recognition method and device based on marking and backtracking correction, and the method comprises the steps: segmenting a target audio, and enabling adjacent segments after segmentation to have a partial overlapping region; if a slice point exists in a non-mute segment of the target audio and the acoustic feature similarity of a preset number of frames before and after the slice point is smaller than a set similarity threshold value, the slice point is marked as a cut-off risk point, and the cut-off risk point is used for indicating that continuous semantics before and after the slice point has a cut-off risk; after a text corresponding to the audio of each slice is recognized through the speech recognition model, the language confidence degree of an overlapping area where the truncation risk point is located is determined, and the language confidence degree is used for indicating context language logic of the overlapping area; and if it is determined that the language confidence is lower than a set confidence threshold, correcting the recognition texts of the slices before and after the truncation risk point. According to the invention, the accuracy of audio recognition is improved.
Owner:FIBOCOM WIRELESS

Broadcast signal interference correction method, system, medium and equipment

The invention discloses a broadcast signal interference correction method and system, a medium and equipment. A broadcast signal is collected in real time and converted into a digital signal, preprocessing is carried out to obtain a first audio signal, feature extraction is carried out according to the first audio signal to obtain audio features, and corresponding preprocessing is carried out on different types of audio features; integrating the audio features to generate an audio feature matrix, inputting the audio feature matrix into a preset fusion model to output interference analysis, receiving the interference analysis, obtaining a corresponding template in combination with a first confidence coefficient, and filling dynamic parameters to generate a correction strategy and the confidence coefficient; wherein the fusion model comprises a decision tree unit and a multi-modal fusion decision unit; and obtaining a correction index corrected by the correction strategy, generating an evaluation index in combination with interference analysis and the correction strategy, and performing adaptive feedback on the fusion model through the evaluation index. Intelligent detection and automatic correction of broadcast signal interference are realized, and the stability and signal quality of a broadcast system are improved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Real-time intention recognition method and system based on streaming incremental reasoning

The invention relates to the technical field of voice processing, in particular to a real-time intention recognition method and system based on streaming incremental reasoning, and the method comprises the following steps: receiving the voice input of a user through a voice collection module, slicing the voice into a plurality of audio frames, and carrying out the recognition of the intention of the user through an incremental large language model module by adopting an Early-Exit reasoning mechanism; side outlets are arranged at multiple levels of the model, incremental reasoning is performed on token streams based on a QLoRA4-bit quantization technology, prediction results of multiple tokens are smoothed by using a stream ASR decoding module and an accumulative fusion module, and a stable final label is generated; the method has the beneficial effects that by combining streaming ASR decoding with incremental large language model reasoning, intention recognition and risk assessment can be immediately performed on each speech token after the speech token is generated. Through an Early-Exit reasoning mechanism, under the condition of high confidence, the system can output a fraud intention in advance in a middle layer in the reasoning process and stop subsequent calculation, and unnecessary calculation overhead is reduced.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Prompt word security reinforcement method and device for large model agent, equipment and medium

According to the embodiment of the invention, the invention provides a cue word security reinforcement method and device for a large model agent, equipment and a medium. The method comprises the following steps: acquiring an initial first system cue word of an intelligent agent based on a large model; determining risk assessment information corresponding to the first system cue word, wherein the risk assessment information indicates at least one risk existing in the first system cue word; and based on the risk assessment information, the first system cue word and the security constraint information, determining a second system cue word of the intelligent agent, so that the intelligent agent processes the received processing request based on the second system cue word. Based on the mode, the security protection of the intelligent agent can be effectively realized, and the confidence and accuracy of the intelligent agent in information processing are improved.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD