An intelligent management method based on real-time speech recognition transcription

By combining multimodal audio perception and deep learning algorithms with semantic error correction and structured processing, the problems of anti-interference and logical association in speech recognition under complex sound field environments are solved, achieving efficient speech transcription and management data conversion, and improving information conversion efficiency and the intelligence level of the management system.

CN122177117APending Publication Date: 2026-06-09MUDANJIANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MUDANJIANG NORMAL UNIV
Filing Date
2026-04-23
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing technologies have insufficient anti-interference capabilities for speech recognition in large-scale multi-turn dialogues or complex sound field environments. The structured processing and core element extraction of transcribed text are not closely related to the actual context logic, resulting in limited efficiency in information sorting during the management process.

Method used

An intelligent management system is constructed by employing multimodal audio perception and environment adaptive enhancement, acoustic feature extraction and real-time transcription mapping, semantic error correction compensation based on dynamic context weighting, structured element extraction and logical association reconstruction, and intelligent management closed-loop decision-making and task distribution, combined with deep learning algorithms and natural language processing technology.

Benefits of technology

It significantly improves the robustness of speech recognition and the accuracy of transcription, realizes efficient conversion from speech signals to structured management data, and improves information conversion efficiency and the intelligence level of management systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177117A_ABST
    Figure CN122177117A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent management method based on real-time speech recognition and transcription, relating to the field of artificial intelligence technology. The method includes the following steps: Step S1, multimodal audio perception and environmental adaptive enhancement; Step S2, acoustic feature extraction and real-time transcription mapping; Step S3, semantic error correction compensation based on dynamic context weighting; Step S4, structured element extraction and logical association reconstruction; Step S5, intelligent management closed-loop decision-making and task distribution; Step S6, multi-source information backtracking and index construction; Step S7, intelligent management efficiency evaluation. This application can solve the problems of limited recognition accuracy and missing semantic logic under complex sound fields. By improving transcription accuracy through multimodal perception and dynamic semantic compensation, it achieves automated connection from speech recording to structured management decisions, significantly improving office management efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an intelligent management method based on real-time speech recognition and transcription. Background Technology

[0002] With the rapid development of artificial intelligence and digital office technology, voice interaction, as the core carrier of information transmission, has received widespread attention for its processing efficiency and accuracy. Real-time speech recognition and transcription, as an important component for realizing intelligent meeting minutes, government office work, and online education, directly determines the level of automation in office management and the convenience of information retrieval through its conversion accuracy, response speed, and semantic understanding capabilities. To ensure the efficient utilization of voice data in multiple scenarios, establishing an intelligent and automated speech transcription management system has become a key research focus in the field of modern information technology.

[0003] Among them, the real-time acquisition, acoustic modeling and language decoding of voice signals mainly rely on deep learning algorithms, automatic speech recognition technology and natural language processing technology. This technical approach aims to achieve rapid conversion of voice content into text information through real-time parsing of audio streams, monitor and record key information in the interaction process, and build digital management capabilities for efficient collaboration.

[0004] Existing technologies have limitations in resisting interference during large-scale multi-turn dialogues or complex sound field environments. Furthermore, the structured processing of transcribed text and the extraction of core elements are not closely related to the actual context, which limits the efficiency of information processing in the management process.

[0005] Therefore, an intelligent management method based on real-time speech recognition and transcription is proposed to solve the above problems. Summary of the Invention

[0006] The main objective of this invention is to provide an intelligent management method based on real-time speech recognition and transcription to solve the problems mentioned in the background above.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: an intelligent management method based on real-time speech recognition and transcription, comprising the following steps: Step S1: Multimodal audio perception and environmental adaptive enhancement; Step S101: Acquire multipath speech signal streams through a pre-distributed microphone array; Step S102: Synchronously call the environmental perception sensor to obtain the background noise energy distribution in the sound field space; Step S103: Perform spatial filtering on the multipath speech signal stream using beamforming algorithm to suppress acoustic interference in non-target areas and obtain the enhanced audio stream to be processed.

[0008] Step S2: Acoustic feature extraction and real-time transcription mapping; Step S201: Input the audio stream to be processed into a preset feature extraction model; Step S202: Extract the spatiotemporal feature vector sequence that reflects the phoneme features of speech; Step S203: Using an end-to-end deep neural network recognition architecture, the feature vector sequence is mapped to the corresponding text candidate set, and a path search is performed in conjunction with a preset acoustic model and dictionary database to generate a preliminary transcribed text stream.

[0009] Step S3: Semantic error correction compensation based on dynamic context weighting; Step S301: Extract contextual features from the preliminary transcribed text stream; Step S302: Substitute the contextual features into the preset semantic evaluation model; Step S303: The semantic evaluation model corrects the confidence of candidate terms by introducing a dynamic context compensation factor. The calculation formula is as follows: ; in, The overall evaluation score for the target text segment. For the first The original recognition probability of each candidate word. This represents the weight coefficient of the term within a specific industry thesaurus. It is a dynamic compensation factor based on the preceding context vector. These are the preset language model smoothing parameters. The length of the current segment being processed.

[0010] Step S4: Extraction of structured elements and reconstruction of logical relationships; Step S401: Perform named entity recognition and dependency parsing on the corrected transcribed text using natural language processing algorithms; Step S402: Extract core management elements such as meeting topic, participants, decision instructions, and execution cycle; Step S403: Construct a logical association matrix based on knowledge graphs to transform fragmented text information into structured management data with temporal logical relationships.

[0011] Step S5: Intelligent management closed-loop decision-making and task distribution; Step S501: Based on the instruction characteristics in the structured management data, automatically match the preset management strategy library, generate a task to-do list, and assign it to the associated terminal execution mechanism. Step S502: By monitoring the task status feedback in real time, a closed-loop mapping between voice recording and management behavior is constructed.

[0012] Furthermore, in step S101, the microphone array is configured as a multi-point distributed arrangement, and each sampling unit maintains phase consistency through a synchronous clock bus; in step S102, the background noise energy distribution is updated in real time by an adaptive noise estimator, and the passband width of the spatial filter is dynamically adjusted by calculating the signal-to-noise ratio; during the signal enhancement process, the Wiener filtering algorithm is used to suppress residual stationary noise, ensuring that the purity of the signal input to the recognition stage is within the preset signal-to-noise ratio range.

[0013] Furthermore, the feature extraction model in step S201 adopts a hybrid architecture of convolutional neural network and recurrent neural network, configured to capture the long-term correlation and local texture features of speech signals; the path search in step S203 adopts a bundle search algorithm with pruning strategy, and the search width is dynamically scaled according to the real-time load status of the computing platform; the dictionary in step S203 includes a preset general vocabulary set and a set of professional terms dynamically loaded according to the current management scenario.

[0014] Furthermore, the dynamic context compensation factor The calculation method is as follows: ; in, This is the semantic vector representation of the current term. This is a global contextual feature vector within a pre-defined historical sliding window. Adjust parameters for environmental complexity. The standard deviation of the semantic distribution is the preset standard deviation. This formula calculates the cosine similarity between the current input and the historical context, and applies weight gain to candidate words that conform to logical continuity.

[0015] Furthermore, S1 also includes a sound source localization and tracking step: using a time difference of arrival algorithm to calculate the spatial coordinates of the target sound source in real time, and driving the directional main lobe of the microphone array to align with the spatial coordinates in real time; if the sound source moving speed exceeds a preset offset threshold, a fast relocation mechanism is triggered, and the directional beam is reconstructed by adjusting the weighting vector weights.

[0016] Furthermore, after S2, multi-criteria streaming output processing is also included: real-time word segmentation and punctuation prediction are performed on the generated preliminary transcribed text stream, and bidirectional long short-term memory network is used to identify sentence boundaries; during the text output process, the decoding delay is adjusted according to the preset real-time level, and a dynamic balance is achieved between recognition accuracy and response speed.

[0017] Furthermore, the named entity recognition in step S401 adopts a coupled model of conditional random field and pre-trained language model, and classifies the key attributes in the text through a preset annotation system; the logical association matrix in step S403 is constructed by calculating the semantic distance and temporal overlap between different elements, and for elements with recognition confidence below a preset threshold, the system automatically marks them as pending confirmation.

[0018] Furthermore, the present invention also includes step S6, multi-source information backtracking and index construction; Step S601: Timestamp-align and store the transcribed text, structured elements, and original audio stream to construct a multidimensional index database; Step S602: The user inputs keywords or logical commands, and the system retrieves related audio segments and management records from the index database and performs a visual presentation.

[0019] Furthermore, the present invention also includes step S7, intelligent management efficiency evaluation; Step S701: Construct a management efficiency evaluation function by collecting response time, completion rate and user feedback data after task distribution; Step S702: Analyze the correlation between voice commands and management results using machine learning algorithms, and automatically optimize the management strategy matching logic in step S501.

[0020] An intelligent management system based on real-time speech recognition and transcription includes a controlled audio acquisition module, a multi-dimensional signal processing unit, an embedded algorithm execution platform, and a management visualization terminal; The controlled audio acquisition module includes an array of microphones and environmental monitoring sensors set up in the office area, with each microphone unit connected to a pre-processing amplifier via a shielded cable; The multidimensional signal processing unit includes a digital signal processor and a high-performance graphics processing core, and integrates the aforementioned acoustic feature extraction model and speech decoding engine. The embedded algorithm execution platform is equipped with a multi-core central processing unit, whose internal logic circuit is divided into a semantic parsing area, a logic reconstruction area and a task scheduling area, and runs the aforementioned error correction and compensation model and structured extraction program. The management visualization terminal is connected to the execution platform through an encrypted communication link. The display interface includes a real-time transcription window, an element extraction dashboard, a task status tracking graph, and a historical record backtracking interface.

[0021] Furthermore, when the embedded algorithm execution platform executes step S4, it maps the extracted entity objects to a preset asset or human resources database according to the ontology definition in the knowledge graph; when executing step S5, if a high-priority instruction is identified, the system prioritizes the processing of the associated task distribution sequence through interrupt control logic.

[0022] Furthermore, the array microphone group in the controlled audio acquisition module adopts MEMS integration technology and has preset dynamic range and frequency response characteristics; the environmental monitoring sensor is configured to monitor the reverberation time and electromagnetic interference level in the space in real time, and feed the parameters back to the multi-dimensional signal processing unit for algorithm parameter calibration.

[0023] The present invention has the following beneficial effects: 1. In this invention, a voice signal stabilization barrier for complex office environments is constructed by integrating multimodal audio perception and environmental adaptive enhancement methods; the spatial filtering capability of the microphone array and noise energy distribution monitoring are used to significantly reduce the erosion of non-target sound sources and environmental background noise on the recognition process; and adaptive noise estimation and beamforming algorithms are used to eliminate the impact of reverberation and abrupt changes in sound field on audio quality, so that the input end of the transcription system has high fidelity characteristics and significantly improves the robustness of subsequent recognition processes.

[0024] 2. In this invention, a high-precision real-time transcription mapping mode is established by physically coupling acoustic feature extraction with a deep neural network architecture; the long-term correlation features of speech signals are captured by an end-to-end recognition architecture, solving the context breakage problem that traditional models are prone to when processing large-scale multi-turn dialogues; and the system achieves significantly improved recognition accuracy for industry-specific terms while maintaining low latency response through dynamic search width adjustment and professional lexicon loading, thereby improving information conversion efficiency in digital office scenarios.

[0025] 3. This invention abandons the traditional static text processing mode and establishes a dynamic semantic compensation mechanism by introducing a semantic error correction model that includes context compensation factors, weight coefficients, and smoothing parameters. This mechanism eliminates semantic deviations caused by homophones or environmental interference by performing mathematical deconvolution operations on the initial transcribed text and comparing it with context vectors, thus significantly enhancing the logical coherence of the transcribed results. Even under the harsh conditions of semantic ambiguity or multiple contexts, it can still maintain a high level of text accuracy, solving the problem of the disconnect between transcribed content and actual context in existing technologies.

[0026] 4. In this invention, by integrating named entity recognition and knowledge graph technologies, deep structural processing and logical association reconstruction are performed on the transcribed text; by utilizing dependency parsing and association matrix construction, a leap from simple text records to management data with decision-making value is achieved; the entire process management closed loop distributes decisions through an embedded platform, realizing the automated connection between voice commands and task execution, significantly improving the intelligence level of the office management system and the convenience of information retrieval, and providing reliable digital support for modern enterprise management. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the overall method architecture of an intelligent management method based on real-time speech recognition and transcription according to the present invention; Figure 2 This is a detailed method architecture diagram of step S1 in the intelligent management method based on real-time speech recognition and transcription of the present invention; Figure 3 This is a detailed method architecture diagram of step S2 in the intelligent management method based on real-time speech recognition and transcription of the present invention; Figure 4 This is a detailed method architecture diagram of step S3 in the intelligent management method based on real-time speech recognition and transcription of the present invention; Figure 5 This is a detailed method architecture diagram of step S4 in the intelligent management method based on real-time speech recognition and transcription of the present invention; Figure 6 This is a detailed method architecture diagram of step S5 in the intelligent management method based on real-time speech recognition and transcription of the present invention; Figure 7 This is a detailed method architecture diagram of step S6 in the intelligent management method based on real-time speech recognition and transcription of the present invention; Figure 8 This is a detailed method architecture diagram of step S7 in the intelligent management method based on real-time speech recognition and transcription of the present invention. Detailed Implementation

[0028] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.

[0029] Please refer to Figures 1 to 8 As shown: An intelligent management method and management system based on real-time speech recognition and transcription is mainly used to achieve high-precision acquisition, transcription and subsequent automated task management of speech information in complex and ever-changing office or meeting sound field environments.

[0030] The system consists of a controlled audio acquisition module, a multi-dimensional signal processing unit, an embedded algorithm execution platform, and a management visualization terminal in terms of physical architecture. The various parts are integrated into an organic technical whole through shielded signal lines, high-speed data buses, and encrypted network protocols.

[0031] In terms of hardware construction, the core components of the controlled audio acquisition module are distributed within the sound field space to be monitored. The module includes an array of microphones arranged in a multi-point distributed manner. Each microphone unit adopts MEMS integration technology and has preset dynamic range and frequency response characteristics. The microphone units are physically connected to the pre-processing amplifier through a synchronous clock bus to ensure the phase consistency of each sampling channel. Environmental sensing sensors are arranged in the gaps between the array of microphones. These sensors are configured to monitor the background noise energy level, reverberation time, and electromagnetic interference level in the space in real time, and convert the monitored analog signals into digital parameters to feed back to the multi-dimensional signal processing unit.

[0032] The multidimensional signal processing unit consists of a digital signal processor and a high-performance graphics processing core. The input of the digital signal processor is connected to the aforementioned array microphone group through a high-speed A / D converter. Its internal logic circuit contains beamforming algorithms and Wiener filtering programs. The high-performance graphics processing core communicates with the digital signal processor through an internal high-speed bus. Its storage area is loaded with a feature extraction model based on a hybrid architecture of convolutional neural networks and recurrent neural networks, as well as a speech decoding engine. The output of this unit is connected to an embedded algorithm execution platform through a parallel interface.

[0033] The embedded algorithm execution platform, serving as the system's decision-making hub, is equipped with a multi-core central processing unit. Its internal logic circuitry is divided into a semantic parsing area, a logic reconstruction area, and a task scheduling area. The semantic parsing area stores preset semantic evaluation models and dynamic context compensation algorithms. The logic reconstruction area accesses a knowledge graph database via an internal bus to perform named entity recognition and dependency parsing. The task scheduling area connects to external terminal execution mechanisms, such as servers in office automation systems or mobile communication terminals, via communication interfaces.

[0034] The management visualization terminal is physically connected to the embedded algorithm execution platform through an encrypted communication link. Its hardware includes a high-resolution display screen and an interactive control panel. The display interface is divided into multiple functional windows, which correspond to real-time transcription stream display, core element dashboard, task status tracking graph, and historical record backtracking interface, respectively.

[0035] Based on the above hardware structure, the specific execution steps of the intelligent management method described in this embodiment are as follows: S1. Multimodal audio perception and environmental adaptive enhancement process: The array microphone group captures the multipath speech signal stream in the sound field, and the background noise energy distribution is obtained simultaneously by the environmental perception sensor; the digital signal processor calls the beamforming algorithm to calculate the spatial coordinates of the target sound source according to the phase difference of each channel signal, and drives the directional main lobe to align with the coordinates in real time to perform spatial filtering; subsequently, the system uses the Wiener filtering algorithm to suppress residual stationary noise, ensuring that the signal-to-noise ratio of the enhanced audio stream is within the preset clean range.

[0036] S2. Acoustic Feature Extraction and Real-time Transcription Mapping Process: The high-performance graphics processing core receives the audio stream to be processed and extracts the spatiotemporal feature vector sequence reflecting the phoneme features of the speech. Using an end-to-end deep neural network recognition architecture, the feature vector sequence is mapped into a text candidate set. At this time, the system combines the preset acoustic model with a dictionary containing general vocabulary and professional terms to perform a bundle search with a pruning strategy to generate a preliminary transcribed text stream.

[0037] S3. Semantic error correction and compensation process based on dynamic context weighting: The embedded algorithm execution platform extracts contextual features from the initial transcribed text stream and substitutes them into the preset semantic evaluation model; the model corrects the confidence of candidate terms by introducing a dynamic context compensation factor, the calculation formula of which is as follows: ; In actual operation, the processor obtains the raw recognition probabilities of candidate words from the semantic vector database. and industry weighting coefficient ; Character length of the currently processed segment Real-time statistics are generated by a counter; language model smoothing parameters. The system adaptively scales based on parameters adjusted according to environmental complexity; among which, a dynamic context compensation factor is used. The calculation method is as follows: ; The system calculates the semantic vector of the current term. Global context feature vector within the historical sliding window The cosine similarity is used to apply weighted gains to text segments that conform to logical continuity, thereby eliminating recognition bias caused by homonyms.

[0038] S4. Structured Element Extraction and Logical Relationship Reconstruction Process: The logical reconstruction area uses natural language processing algorithms to perform deep analysis on the corrected text; it identifies core elements such as meeting topics, participants, and decision instructions through named entity recognition, and uses dependency parsing to determine the modification and dominance relationships between elements; subsequently, the system constructs a logical relationship matrix, transforming fragmented text into structured management data with temporal logical relationships. For elements with identification confidence levels below the threshold, they are automatically marked as pending confirmation on the visualization terminal.

[0039] S5. Intelligent management closed-loop decision-making and task distribution process: The task scheduling area automatically matches the management plan in the strategy library according to the instruction characteristics in the structured management data and generates a task list; the list is distributed to the associated terminal execution agencies in real time through the network interface; the system continuously monitors the status feedback of each execution terminal and sends the feedback information back to the visualization terminal, forming a closed-loop mapping between voice recording and management behavior.

[0040] Step S6, Multi-source Information Backtracking and Index Construction: The transcribed text, structured elements, and original audio stream are timestamped and stored to construct a multi-dimensional index database; users can input keywords or logical commands, and the system can retrieve related audio segments and management records from the index database and perform visualization presentation.

[0041] Step S7, Intelligent Management Efficiency Evaluation: By collecting response time, completion rate and user feedback data after task distribution, a management efficiency evaluation function is constructed; machine learning algorithms are used to analyze the correlation efficiency between voice commands and management results, and the management strategy matching logic in step S5 is automatically optimized.

[0042] In specific application scenarios, such as cross-departmental coordination meetings in large enterprises, the system is deployed in the central area of ​​the conference room. When multiple participants take turns speaking, the microphone array drives the directional main lobe to switch quickly between different speakers through a sound source localization and tracking mechanism. Even under background noise interference such as air conditioning operation or paper turning, the multi-dimensional signal processing unit can still maintain stable audio quality through spatial filtering. When processing the transcription stream, the embedded algorithm execution platform automatically identifies key instructions such as "complete the financial statement audit by next Wednesday" and uses a dynamic context compensation model to ensure the accuracy of professional terms such as auditing and finance. Subsequently, the system automatically extracts the task executor and deadline, converts them into a to-do process in the OA system, and displays them on the management dashboard in real time.

[0043] The system's physical connections not only ensure the sensitivity of signal acquisition but also guarantee the immediacy of management commands. The shielded cable between the array microphone group and the preprocessing amplifier effectively resists electromagnetic pulse interference in the office environment. The embedded algorithm execution platform interacts with the multi-dimensional signal processing unit via a high-speed bus, and its internal non-volatile memory chip records professional lexicon models from different industries. When the management scenario changes from an administrative meeting to a technical discussion, the operator only needs to select the corresponding industry code on the management visualization terminal, and the system can automatically load the corresponding relative correction factor and semantic distribution standard deviation database. Through this highly integrated hardware layout and dynamic correction algorithm, this embodiment effectively solves the management deviation caused by sound field interference and semantic logic deficiency, and significantly improves the application accuracy of speech transcription in intelligent office.

[0044] To enable those skilled in the art to fully understand and implement this invention, the specific implementation principles of this invention are further supplemented below with a specific application scenario.

[0045] In the context of digital government office work, when it is necessary to intelligently manage government meetings that contain a large number of legal terms and multi-party interactive instructions, the specific operational principle is as follows: Step 1: Audio Stabilization and Sound Field Reconstruction Operating Principles: First, the acoustic characteristics of the conference room are monitored in real time by environmental perception sensors. When the reverberation time exceeds a set threshold, the embedded algorithm execution platform drives the digital signal processor to adjust the passband width of the spatial filter and uses beamforming technology to construct a virtual pickup area in the physical space. During this process, the microphone array calculates the arrival time difference and locks the directional main lobe at the geometric coordinate point of the current speaker. It uses the principle of spatial phase cancellation to suppress noise interference in non-target areas, ensuring that the audio stream entering the feature extraction stage has high signal-to-noise ratio characteristics.

[0046] Step 2, Multidimensional Semantic Chain Synchronous Parsing and Enhancement Principle: During the transcription task, the audio stream is transformed into a high-dimensional vector through the feature extraction model; the high-performance graphics processing kernel uses a deep neural network to decode the vector sequence in parallel; at the same time, the embedded algorithm execution platform sends a context synchronization signal to the semantic parsing area, enabling candidate words in the search space to be dynamically sorted according to the current government context; this physical-level algorithm linkage enables the system to identify specific government terms that are spoken quickly or have a regional accent, eliminating the risk of missed identification caused by the insufficient generalization ability of a single acoustic model.

[0047] Step 3: Logical Element Mapping and Closed-Loop Management Principle: In the logical reconstruction phase, the system retrieves the logical association matrix from step S4 for calculation. At this time, the processor searches the ontology definition in the knowledge graph in real time through the internal bus, mapping the identified participants to the job responsibilities in the human resources database. Due to the strict temporal nature of government instructions, the processor quantitatively assesses the urgency of management tasks by calculating the semantic distance between the instruction verbs and time entities. Furthermore, by monitoring the online status of task execution terminals, the distribution route in step S5 is automatically optimized. This real-time connection based on semantic logic and physical execution status ensures that the final generated task list accurately corresponds to the responsible entity, restoring the true business logic in the office management process.

[0048] Step 4: Closed-Loop Feedback Execution and Efficiency Optimization Principle: When the management visualization terminal receives a feedback signal indicating task completion, the task scheduling area of ​​the embedded algorithm execution platform immediately enters the update state. If a task response time triggers an early warning, the processor outputs a control signal through logic circuits, driving the system to automatically send reminders or adjust subsequent task distribution strategies. This achieves a physical closed loop from audio acquisition and high-precision transcription to automated management feedback. Through precise manipulation of management behavior by algorithms, the operational efficiency of the office system is significantly improved.

[0049] To enable those skilled in the art to fully understand and implement this invention, the specific implementation principles of this invention are further supplemented below with a specific application scenario.

[0050] Step 1: Sound Field Perception and Signal Purification Operating Principle: In the application scenario of a large tiered conference room, the array microphone group in the controlled audio acquisition module is distributed at preset coordinate points on the stage and in the audience seating area. When the speaker begins to speak, the environmental monitoring sensor collects the reverberation parameters and background noise energy levels in the sound field in real time, and converts these physical quantities into electrical signals that are fed back to the multi-dimensional signal processing unit. The digital signal processor calculates the speaker's spatial pointing vector based on the arrival time difference of the signals received by each microphone unit through its internal beamforming logic. At this time, the system drives the directional main lobe to align with this vector coordinate, and uses the principle of spatial phase cancellation to physically attenuate the sound wave energy in the non-directional area. Subsequently, the Wiener filtering program performs spectral subtraction on the audio stream to be processed based on the noise benchmark provided by the environmental monitoring sensor, eliminating stable noises such as air conditioning fan noise, thereby completing the extraction and purification of speech at the physical signal level.

[0051] Step Two: Feature Mapping and Candidate Stream Generation Operation Principle: The high-performance graphics processing core receives the purified audio stream and uses convolutional neural network logic to segment the continuous audio waveform into a frame sequence of preset length, extracting the spatiotemporal feature vector of each frame. These vector sequences are input into an end-to-end deep neural network recognition architecture. The model uses internally stored acoustic distribution probabilities to map the feature vectors to corresponding text candidate sets. During path search, the system calls a dictionary containing domain-specific vocabulary and, based on the pruning probability threshold of the beam search algorithm, eliminates low-confidence character combinations. At this point, the system dynamically adjusts the search width according to the real-time load of the computing platform, completing the initial encapsulation and transmission of the transcribed text stream within the memory bus bandwidth.

[0052] Step 3: Semantic Logic Correction and Context Compensation. Operating Principle: The embedded algorithm execution platform receives the initial transcribed text stream and loads it into the semantic parsing area. The processor extracts the semantic vector of the current text segment. It also retrieves the global contextual feature vector within the preceding sliding window from non-volatile memory. System execution formula:

[0053] ; Among them, dynamic context compensation factor Through calculation and The cosine similarity is used to obtain the result; when homophones such as "audit" or "record" are identified, if the current context vector has a very high similarity to the "finance" related vector in the historical context, then... This will assign a higher weight gain to "audit"; language model smoothing parameters Adaptive scaling is performed based on complexity parameters fed back from environmental monitoring sensors, thereby correcting recognition biases at the mathematical level and ensuring that the text stream conforms to logical continuity.

[0054] Step 4: Structured Management and Closed-Loop Task Operation Principle: The logic reconstruction zone uses a conditional random field model to perform deep analysis on the corrected text, identifying core elements such as attendees, time entities, and action instructions. The system uses dependency parsing to determine "who" performs "what task" "when" and constructs a logical association matrix accordingly. Subsequently, the task scheduling zone maps the elements in the logical association matrix to a preset management strategy library. If the instruction characteristics trigger a high-priority threshold, the processor outputs an interrupt control signal, driving the communication interface to instantly distribute the generated task list to the associated office automation server. The management visualization terminal retrieves the task status tracking graph in real time and aligns the physical conversion process of voice instructions with the task execution status by monitoring the feedback signals from the execution terminal, completing the closed-loop control from sound wave acquisition to management behavior feedback.

[0055] All contents not described in detail in the specification are existing technologies known to those skilled in the art, and the model parameters of each electrical appliance are not specifically limited; conventional equipment can be used. Electrical control components not mentioned in this technical solution are not shown in the figures because they are existing technologies, and will not be described here.

[0056] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent management method based on real-time speech recognition and transcription, characterized in that, Includes the following steps: Step S1: Multimodal audio perception and environmental adaptive enhancement; Step S101: Acquire multipath speech signal streams through a pre-distributed microphone array; Step S102: Synchronously call the environmental perception sensor to obtain the background noise energy distribution in the sound field space; Step S103: Perform spatial filtering on the multipath speech signal stream using beamforming algorithm to suppress acoustic interference in non-target areas and obtain the enhanced audio stream to be processed; Step S2: Acoustic feature extraction and real-time transcription mapping; Step S201: Input the audio stream to be processed into a preset feature extraction model; Step S202: Extract the spatiotemporal feature vector sequence that reflects the phoneme features of speech; Step S203: Using an end-to-end deep neural network recognition architecture, the feature vector sequence is mapped to the corresponding text candidate set, and a path search is performed in combination with a preset acoustic model and dictionary database to generate a preliminary transcribed text stream. Step S3: Semantic error correction compensation based on dynamic context weighting; Step S301: Extract contextual features from the preliminary transcribed text stream; Step S302: Substitute the contextual features into the preset semantic evaluation model; Step S303: The semantic evaluation model corrects the confidence of candidate terms by introducing a dynamic context compensation factor. Step S4: Extraction of structured elements and reconstruction of logical relationships; Step S5: Intelligent management closed-loop decision-making and task distribution.

2. The intelligent management method based on real-time speech recognition and transcription according to claim 1, characterized in that, The microphone array in step S101 is configured as a multi-point distributed arrangement, and each sampling unit maintains phase consistency through a synchronous clock bus; the background noise energy distribution in step S101 is updated in real time by an adaptive noise estimator, and the passband width of the spatial filter is dynamically adjusted by calculating the signal-to-noise ratio; during the signal enhancement process, the Wiener filtering algorithm is used to suppress residual stationary noise, so that the purity of the signal input to the recognition stage is within the preset signal-to-noise ratio range.

3. The intelligent management method based on real-time speech recognition and transcription according to claim 1, characterized in that, Step S1 also includes a sound source localization and tracking step: the spatial coordinates of the target sound source are calculated in real time using a time difference of arrival algorithm, and the directional main lobe of the microphone array is driven to align with the spatial coordinates in real time; if the sound source moving speed exceeds a preset offset threshold, a fast relocation mechanism is triggered, and the directional beam is reconstructed by adjusting the weighting vector weights.

4. The intelligent management method based on real-time speech recognition and transcription according to claim 1, characterized in that, The feature extraction model in step S201 adopts a hybrid architecture of convolutional neural network and recurrent neural network, configured to capture the long-term correlation and local texture features of speech signals; the path search in step S203 adopts a bundle search algorithm with pruning strategy, and the search width is dynamically scaled according to the real-time load status of the computing platform; the dictionary in step S203 includes a preset general vocabulary set and a professional terminology set dynamically loaded according to the current management scenario.

5. The intelligent management method based on real-time speech recognition and transcription according to claim 1, characterized in that, The formula for correcting the calculation in step S3 is as follows: ; in, The overall evaluation score for the target text segment. For the first The original recognition probability of each candidate word. This represents the weight coefficient of the term within a specific industry thesaurus. It is a dynamic compensation factor based on the preceding context vector. These are the preset language model smoothing parameters. The length of the current segment being processed.

6. The intelligent management method based on real-time speech recognition and transcription according to claim 5, characterized in that, The dynamic context compensation factor The calculation method is as follows: ; in, This is the semantic vector representation of the current term. This is a global contextual feature vector within a pre-defined historical sliding window. Adjust parameters for environmental complexity. The standard deviation of the semantic distribution is the preset value.

7. The intelligent management method based on real-time speech recognition and transcription according to claim 1, characterized in that, Step S4 specifically includes the following steps: Step S401: Perform named entity recognition and dependency parsing on the corrected transcribed text using natural language processing algorithms; Step S402: Extract the core management elements of the meeting topic, participants, decision-making instructions, and execution cycle; Step S403: Construct a logical association matrix based on knowledge graphs to transform text information into structured management data with temporal logical relationships; The named entity recognition in step S401 adopts a coupled model of conditional random field and pre-trained language model, and classifies the key attributes in the text through a preset annotation system; the logical association matrix in step S403 is constructed by calculating the semantic distance and temporal overlap between different elements. For elements with recognition confidence below a preset threshold, the system automatically marks them as pending confirmation.

8. The intelligent management method based on real-time speech recognition and transcription according to claim 1, characterized in that, Step S5 specifically includes the following steps: Step S501: Based on the instruction characteristics in the structured management data, automatically match the preset management strategy library, generate a task to-do list, and assign it to the associated terminal execution mechanism. Step S502: By monitoring the task status feedback in real time, a closed-loop mapping between voice recording and management behavior is constructed.

9. The intelligent management method based on real-time speech recognition and transcription according to claim 1, characterized in that, Also includes: Step S6: Multi-source information backtracking and index construction; Step S601: Timestamp-align and store the transcribed text, structured elements, and original audio stream to construct a multidimensional index database; Step S602: The user inputs keywords or logical commands, and the system retrieves related audio segments and management records from the index database and performs a visual presentation. Step S7: Evaluation of the effectiveness of intelligent management; Step S701: Construct a management efficiency evaluation function by collecting response time, completion rate and user feedback data after task distribution; Step S702: Analyze the correlation between voice commands and management results using machine learning algorithms, and automatically optimize the management strategy matching logic in step S501.

10. An intelligent management system based on real-time speech recognition and transcription, referring to the intelligent management method based on real-time speech recognition and transcription according to any one of claims 1-9, characterized in that, It includes a controlled audio acquisition module, a multi-dimensional signal processing unit, an embedded algorithm execution platform, and a management visualization terminal; The controlled audio acquisition module includes an array of microphones and environmental monitoring sensors set up in the office area, with each microphone unit connected to a pre-processing amplifier via a shielded cable; The multidimensional signal processing unit includes a digital signal processor and a high-performance graphics processing core, and integrates an acoustic feature extraction model and a speech decoding engine. The embedded algorithm execution platform is equipped with a multi-core central processing unit, whose internal logic circuit is divided into a semantic parsing area, a logic reconstruction area and a task scheduling area, and runs an error correction and compensation model and a structured extraction program. The management visualization terminal is connected to the execution platform through an encrypted communication link. The display interface includes a real-time transcription window, an element extraction dashboard, a task status tracking graph, and a historical record backtracking interface.