A speech analysis method, system, device and medium for multi-task parallel processing
Through the speech analysis method of multi-task parallel processing, multiple speech tasks are generated and processed in parallel, and the speaker information flow is allocated and analyzed according to priority. The problem of long waiting time for existing systems is solved, and the speaker's rapid response and eloquence performance in complex environments is improved.
Patent Information
- Application Number
- CN202410880338.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-02
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-07-02
AI Technical Summary
The existing eloquence training system cannot simulate multiple problems and raise them at the same time, resulting in the waiting time of the analysis process, which cannot output analysis results to the speakers in time, and cannot effectively improve the speakers' ability to think quickly and react when facing emergencies and complex problems.
Through the speech analysis method of multi-task parallel processing, multiple speech tasks are generated, and the speakers are assigned simultaneously according to priority. The information flow generated by the speakers is processed in parallel, and the eloquence performance analysis is performed to generate expression improvement suggestions.
It improves the speaker's ability to think quickly and react when facing emergencies and complex problems, improves the efficiency of training, and enhances the speaker's eloquence and expression.
Smart Images

Figure CN118657157B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of oratory training data analysis, and in particular to a speech analysis method, system, device and medium for multi-task parallel processing. Background Art
[0002] In order to improve the quick thinking and response ability of speakers when facing emergencies and complex problems, a Q&A session is often added during the speaker's speech, that is, relevant questions are raised by on-site audiences or teachers for the speaker to answer. At present, in order to improve the professionalism of training, the existing developed oratory expression training systems can only simulate that the audience or teacher single-handedly asks a single question to the speaker, and the system will ask another question after the speaker finishes answering the question. The existing oratory training systems cannot simulate the high-difficulty scenario of multiple questions being asked simultaneously; and when the system analyzes the answers to the questions answered by the speaker, it also needs to analyze the answer to one question before analyzing the answer to another question, which will lead to a long waiting process during the analysis and cannot output the corresponding analysis results to the speaker in a timely manner. Summary of the Invention
[0003] Embodiments of the present invention provide a speech analysis method, system, device and medium for multi-task parallel processing to solve the problems existing in the related technologies. The technical solutions are as follows:
[0004] In a first aspect, an embodiment of the present invention provides a speech analysis method for multi-task parallel processing, including:
[0005] Obtain speech content and generate multiple speech tasks according to the speech content;
[0006] Determine multiple target tasks to be allocated according to the priorities of all speech tasks, and allocate the multiple target tasks to the speaker simultaneously at a specified time point;
[0007] Obtain the information flow generated when the speaker processes multiple target tasks in parallel, perform oratory performance analysis on each information flow, and generate corresponding expression improvement suggestions.
[0008] In an implementation manner, the method for determining the priority is:
[0009] Identify the speech content to obtain multiple information units;
[0010] Count the number of information units associated with each speech task;
[0011] Evaluate the information density contained in each information unit to obtain the load of the information unit;
[0012] Calculate the information load corresponding to each speech task according to the number of information units and the load of the information unit, and determine the priority according to the information load.
[0013] In one embodiment, it further includes:
[0014] When obtaining the information flow generated when the speaker processes multiple target tasks in parallel, identify the key features of each information flow, find the corresponding target tasks for each information flow according to the key features, and bind the information flow to the corresponding target task.
[0015] In one embodiment, conduct an oratory performance analysis on each information flow, including:
[0016] Extract the index features associated with the preset index in the information flow;
[0017] Calculate the index scores of each information flow on the preset index according to the index features.
[0018] In one embodiment, the preset index includes topic coherence, emotional engagement, information density, audience interactivity, and logic.
[0019] In one embodiment, generate corresponding expression improvement suggestions, including:
[0020] Obtain the speaking goals corresponding to each target task;
[0021] Judge whether the index score corresponding to each target task reaches its corresponding speaking goal. If the index score of any target task does not reach its corresponding speaking goal, generate corresponding expression improvement suggestions according to the index score.
[0022] In one embodiment, it further includes:
[0023] Dynamically adjust the number of target tasks assigned to the speaker according to the index score.
[0024] In a second aspect, an embodiment of the present invention provides a speech analysis system that executes the speech analysis method for multi-task parallel processing as described above.
[0025] In a third aspect, an embodiment of the present invention provides an electronic device, which includes: a memory and a processor. Among them, the memory and the processor communicate with each other through an internal connection path. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory. When the processor executes the instructions stored in the memory, the processor executes the method in any one of the above aspects.
[0026] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a computer, the method in any one of the above aspects is executed.
[0027] The advantages or beneficial effects in the above technical solutions at least include:
[0028] The present invention can generate multiple speech tasks according to the speech content, and simultaneously assign multiple target tasks to the speaker according to the priority. After receiving the multiple target tasks, the speaker needs to respond to each target task to generate corresponding information flows. By performing parallel processing and analysis on the information flows corresponding to each target task, the rapid thinking and reaction ability of the speaker when facing sudden situations and complex problems can be evaluated, and thus corresponding oratory expression suggestions can be generated. During this process, multiple target tasks are simultaneously assigned to the speaker, which can meet the training needs of the speaker in high-difficulty scenarios, and perform parallel processing and analysis on the information flows generated by the speaker for parallel processing of multiple target tasks, solving the problem of too long analysis waiting time caused by the traditional system queuing analysis and improving the processing efficiency.
[0029] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present invention will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in accordance with the present invention and should not be regarded as limiting the scope of the present invention.
[0031] Figure 1 is a schematic flowchart of the speech analysis method for multi-task parallel processing of the present invention;
[0032] Figure 2 is a block diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] In the following, only some exemplary embodiments are briefly described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present invention. Therefore, the drawings and the description are considered to be exemplary in nature and not restrictive.
[0034] In order to improve the speaker's ability to think quickly and respond when facing emergencies and complex problems, a Q&A session is often added during the speaker's speech, where relevant questions are raised by the on-site audience or teachers for the speaker to answer. Currently, in order to improve the professionalism of training, the existing developed eloquence expression training systems can simulate the audience or teachers to automatically ask questions to the speaker. However, all existing eloquence expression training systems adopt a one-on-one Q&A form, that is, one question is raised, and after the speaker completes the answer to that question, another question is raised. The speaker only needs to think about the answering ideas for a single question during the Q&A session, and its difficulty is relatively weak, and the effect of exercising the speaker's mental agility and response ability is not good.
[0035] To solve the above problems, this embodiment provides a speech analysis method for multi-task parallel processing, which can allocate multiple tasks to the speaker and train the speaker to execute multiple tasks simultaneously to improve the speaker's ability to think quickly and respond when facing emergencies and complex problems. As Figure 1 shown, the speech analysis method includes the following steps:
[0036] Step S1: Obtain the speech content and generate multiple speech tasks according to the speech content;
[0037] Step S2: Determine multiple target tasks to be allocated according to the priorities of all speech tasks, and allocate the multiple target tasks to the speaker simultaneously at a specified time point;
[0038] Step S3: Obtain the information flow generated when the speaker processes multiple target tasks in parallel, perform an eloquence performance analysis on each information flow, and generate corresponding expression improvement suggestions.
[0039] Among them, the speech content can be the speech text input by the speaker according to his speech needs, and the speech ideas, speech themes, etc. of this speech can be recorded in the speech text, or it can also be a complete speech manuscript.
[0040] Preprocess the speech content, including segmenting / sentencing the speech content through natural language analysis algorithms, extracting keywords for each segment / sentence through natural language analysis algorithms, analyzing the core purpose of each segment / sentence through AI algorithms, etc., and dividing multiple speech tasks according to the analyzed core purpose; among them, natural language analysis algorithms and AI algorithms are both existing technologies and will not be described in detail here.
[0041] For example, assume that a speaker is preparing a speech on "sustainable development", and the speech content includes the following parts:
[0042] - Introduction: Introduce the theme and importance.
[0043] - Current situation analysis: Provide data and facts on current environmental problems.
[0044] - Solution: Propose solutions to environmental problems.
[0045] - Call to action: Encourage the audience to participate and support.
[0046] The speech tasks that can be identified based on the speech content include the introduction task, the current situation analysis task, the solution task, and the action task.
[0047] In a speech, there must be some parts that are the key points of the speech and some parts that are secondary. Therefore, in this embodiment, a corresponding priority is set for each task, and the method for determining the priority is as follows:
[0048] Identify the speech content to obtain multiple information units;
[0049] Count the number of information units associated with each speech task;
[0050] Evaluate the information density contained in each information unit to obtain the load of the information unit;
[0051] Calculate the information load corresponding to each speech task based on the number of information units and the load of the information unit, and determine the priority according to the information load.
[0052] Among them, an information unit refers to the smallest meaningful unit that constitutes the speech content, and they are the basic components of speech information transmission; an information unit can be a word, a phrase, a sentence, or a paragraph composed of several sentences, depending on the complexity of the speech content and the information granularity that the speaker wants to convey. The methods for determining information units include:
[0053] 1. Text analysis:
[0054] Through natural language processing (NLP) technology, perform text analysis on the speech content to identify semantically relatively independent basic units.
[0055] 2. Semantic segmentation:
[0056] Use a sentence boundary detection algorithm to determine sentence-level information units, or use a finer-grained analysis to determine word- or phrase-level information units.
[0057] 3. Content structure:
[0058] Through natural language processing (NLP) technology, analyze the structure of the speech content, such as the introduction, the body, and the conclusion, and further divide each part into more specific information units.
[0059] 4. Logical relationship:
[0060] Identify the logical relationships in a speech, such as causality, transition, parallelism, etc., through natural language processing (NLP) techniques to determine the boundaries of information units.
[0061] Divide the speech content into multiple information units, count the number of information units associated with each speech task, evaluate the information density contained in each information unit, and obtain the load of the information unit; among them, the calculation of information density is through the method of information theory, such as calculating the information entropy of each information unit in the speech to evaluate the information density; or identify and quantify the number of key information points to evaluate the information density, and the magnitude of this information density can reflect the load of the information unit.
[0062] In order to scientifically and accurately evaluate the priority of speech tasks, a mathematical model can be established to calculate the information load, and then determine the task priority. This mathematical model should consider two factors, the number and load of information units, because these two parameters can reflect the complexity and information density of speech tasks. Specifically, the following formula can be used to calculate the information load of each speech task:
[0063]
[0064] Among them, L task represents the information load of the task, N t is the total number of information units contained in task t, N i is the number of information unit i, and I i is the load of the i-th information unit. The load of the information unit can be quantified by information entropy or the number of key information points, reflecting the amount of information or importance carried by the information unit.
[0065] The information load is associated with the priority. The higher the priority, the more important the task is and it should be allocated to the speaker first. The calculation of the priority can consider the information load and the estimated processing time of the task, and the formula is as follows:
[0066]
[0067] Here, P task represents the priority of the task, L task represents the information load of the task, T t is the estimated processing time of task t. Through this formula, tasks with high information density and short estimated processing time will be assigned higher priorities.
[0068] In this embodiment, the number of tasks for parallel processing is determined according to the requirements of the speaker, and then the speech tasks to be assigned to the speaker and the quantity are determined by combining the priority levels of all speech tasks. The specified number of speech tasks to be assigned are all marked as target tasks. The target tasks are simultaneously pushed to the speaker in a specified form at a specified time point. The specified form can be provided to the speaker in the form of voice and / or text, enabling the speaker to perform parallel processing on all target tasks. The specified time point can be when the speaker actively triggers the task assignment function to simultaneously push multiple target tasks to the speaker, or it can be a preset time node, such as automatically assigning target tasks after the speaker's speech ends.
[0069] When the speaker receives multiple target tasks, they are executed simultaneously; the voice data, expression data, body data, etc. generated when the speaker executes multiple target tasks are collected as information streams, and each information stream is analyzed.
[0070] In the case of obtaining the information streams generated when the speaker performs parallel processing on multiple target tasks, since the viewpoints expressed by the speaker are not directed at a single target task, it is necessary to identify the key features of the voice data through natural language analysis algorithms or AI algorithms, and find the target tasks associated with them through the semantics of the key features, so as to bind each information stream to its corresponding target task and avoid the situation where the question and answer content do not match. At the same time, the voice data, expression data, and body data are synchronized on the time axis. If it is recognized that the voice data in a certain time period is associated with a certain target task, the corresponding expression data and body data in this time period are found according to the time axis of the voice data and associated with the corresponding target task. The method of binding each information stream to its corresponding target task is as follows:
[0071] Information stream feature extraction: First, deeply analyze the information stream and extract key features. This can be achieved through natural language processing technology (NLP), identifying keywords, semantic units, and emotional tendencies in the voice data. For example, use deep learning models such as recurrent neural networks (RNN) or attention mechanisms to capture semantic and emotional cues in the information stream.
[0072] Task feature mapping: Second, establish a task feature library and encode the key information of each task, such as task theme, expected emotional expression, logical structure, etc. This helps with subsequent information stream and task matching.
[0073] Semantic Relevance Calculation: Use semantic similarity algorithms (such as cosine similarity, Jaccard similarity) to compare the relevance of information flow features with each task in the task feature library and find the most relevant task. For non-semantic indicators such as emotional investment and logic, sentiment analysis and logic analysis algorithms can be used to evaluate the fit between the information flow and the task in these aspects.
[0074] Timeline Synchronization and Cross-Validation: Considering that the information flow may contain responses to multiple tasks, it is necessary to synchronize the timelines of voice data, facial expression data, and body data, and determine which data segments coincide with the time window of a specific task through time series analysis. For example, if the voice data within a certain time period overlaps with the question time of Task A, then this voice data is very likely a response to Task A.
[0075] Deep Learning Model Assistance: Use deep learning models to enhance the accuracy of feature recognition and binding, especially for those subtle features and implicit semantic associations. The model can be pre-trained or fine-tuned on a large amount of speech data to improve the binding accuracy.
[0076] Real-Time Feedback and Correction: Introduce a real-time feedback mechanism, that is, while analyzing the information flow, dynamically adjust the task allocation or prompt the speaker to improve according to the analysis results. This immediate feedback can enhance the immediacy and effectiveness of training.
[0077] Prevention of Incorrect Binding Errors: Reduce incorrect bindings by setting thresholds and multi-layer verification mechanisms. For example, only when the relevance of the information flow to a certain task exceeds a certain threshold will it be considered a correct binding. In addition, a verification stage can be set up to confirm the accuracy of the binding using manual review or more complex algorithms.
[0078] When identifying and binding the information flow, an advanced deep learning model is adopted to improve the accuracy of identification and reduce incorrect binding errors. The model is trained with a large number of samples and can capture the implicit meaning between the subtle features in the information flow and the tasks, ensuring the accuracy of the analysis. In addition, a real-time feedback mechanism is introduced, that is, the information flow analysis is synchronized with the task adjustment to ensure timely feedback and guidance, enhancing the immediacy of training.
[0079] This embodiment mainly uses the voice data of the speaker's oral expression as the main data source for the analysis of oratory expression, analyzes the index scores of oratory expression under each preset index, and thus judges the speaker's quick thinking and reaction ability in the case of multiple target tasks processed in parallel.
[0080] Among them, the methods for analyzing oratory performance include:
[0081] Extract the metric features associated with the preset metrics in the information flow; where the preset metrics include topic coherence, emotional engagement, information density, audience interactivity, and logicality; calculate the metric scores of each information flow on the preset metrics according to the metric features.
[0082] When the preset metric is topic coherence, the semantic relevance between contexts in each information flow can be calculated through text analysis, for example, calculating the similarity between sentences or paragraphs through cosine similarity or Jaccard similarity, and at the same time, it can also be analyzed whether the topic corresponding to the information flow is consistent with the topic of the target task. In the actual analysis process, the metric score of the topic coherence of the information flow is calculated according to the context similarity and the topic consistency procedure.
[0083] When the preset metric is emotional engagement, the emotional color expressed by the information flow is evaluated based on a pre-trained emotion dictionary or a machine learning model. For example, analyzing the tone and intensity of speech, as well as non-verbal signals such as facial expressions and body language, to evaluate the degree of emotional expression, so as to obtain the metric score of emotional engagement.
[0084] When the preset metric is information density, the density of information is evaluated through methods of information theory, such as calculating the information entropy of each unit in a speech; or identifying and quantifying the number of key information points and their distribution in the speech to evaluate the information density, and obtaining the corresponding metric score according to the information density situation.
[0085] If there is audience feedback data in the collected information flow, the metric score under the audience interactivity metric can be calculated for the audience feedback data; if there is no audience feedback data collected in the information flow, there is no need to calculate the metric score under the audience interactivity metric. When the preset metric is audience interactivity, the interaction frequency is evaluated by analyzing the audience feedback and the speaker's response in the Q&A session of the speech, or using social network analysis techniques to measure the audience participation, evaluating the interaction frequency between the speaker and the audience, and obtaining the metric score under the audience interactivity metric according to the interaction frequency.
[0086] When the preset metric is logical rigor, evaluate the structure, validity, and consistency of the arguments in the speech content to obtain the metric score under the logicality metric. Specifically:
[0087] Argument structure identification: Natural language processing techniques can be used to analyze the text content in the speech information flow to identify basic argument elements such as arguments (i.e., the views or claims put forward by the speaker), evidence (facts, data, examples supporting the argument), and rebuttals (responses to potential objections). This step may involve parsing the relationships between sentences, such as identifying logical connectives such as causality, conditionality, and contrast, to ensure an accurate grasp of the argument framework.
[0088] For validity assessment, through AI technology, following the methods of formal logic, it is possible to check whether the formal structure of the argument conforms to logical rules. For example, ensure that each argument follows a valid reasoning pattern (such as the syllogism in deductive logic). In addition, evaluate whether the evidence is directly relevant and fully supports the claim, and whether there are logical fallacies (such as reversing causation, irrelevant argument, etc.). Informal logic analysis focuses on the substantive content of the argument, including the reliability of the evidence, the practical applicability of the claim, and the reasonableness of the refutation. This involves a qualitative assessment of the evidence, such as whether the evidence is reliable, whether the source is authoritative, and whether the refutation is strong enough to effectively address potential challenges.
[0089] Logical consistency checking is used to ensure that all claims, evidence, and refutations in a speech are logically consistent and there are no self - contradictions. This includes checking whether different parts of the argument logically support each other without logical conflicts, and whether the logical chain of the entire speech is coherent and unbroken.
[0090] Based on the above analysis, a score is given for logical rigor. This usually involves quantitative analysis. For example, according to factors such as the number of logical errors found, the clarity of the argument, the sufficiency of the supporting materials, and the effectiveness of the refutation, a set of scoring criteria or weight systems are pre - established.
[0091] The scoring can be in a five - point system, a hundred - point system, or other grades. The specific scoring criteria may vary according to the goal of the speech, the expected level of the audience, and specific training requirements. For example, a speech with strict logic will receive a high score due to its clear argument structure, sufficient evidence, and logical consistency.
[0092] The process of scoring logical rigor can be a process of deeply analyzing the logical content of the speech based on natural language analysis methods, AI algorithms, or manual analysis and other techniques. By identifying the argument structure, evaluating the validity and consistency of the argument, and finally giving a score according to a series of logical quality criteria, it aims to reflect the logical clarity and accuracy of the speaker when constructing and conveying complex ideas.
[0093] After obtaining the index scores corresponding to each index in this embodiment, the comprehensive oratory index score can be calculated according to the index scores and the corresponding weights of each index score, which is used to evaluate the performance of the speaker on multiple advanced oratory indexes. The formula is:
[0094]
[0095] Among them, Uk, Vk, Dk, Hk, and Mk respectively represent the index scores of the k - th target task in terms of topic coherence, emotional engagement, information density, audience interactivity, and logical rigor, and γ, δ, and θ are the corresponding weight indices.
[0096] Generate corresponding expression improvement suggestions according to the index scores, and the corresponding method includes the following steps:
[0097] Obtain the speech goals corresponding to each target task, and the speech goals can be the preset standard score ranges for each index;
[0098] Judge whether the index score corresponding to each target task reaches its corresponding speech goal. In the case that the index score of any target task does not reach its corresponding speech goal, generate corresponding expression improvement suggestions according to the index score.
[0099] Alternatively, according to the comprehensive oratory index score, search for the corresponding expression improvement suggestions under this comprehensive oratory index score in the preset database.
[0100] Meanwhile, the difficulty of parallel processing can also be dynamically adjusted according to each index score. For example, when each index score is relatively low, the number of target tasks for the speaker to process in parallel can be reduced. If each index score meets the standard, the number of tasks for parallel processing can be increased.
[0101] The specific adjustment rules can be set as follows:
[0102] 1. When all index scores reach or exceed their respective preset baseline values, the system can increase the number of tasks for the speaker to process in parallel. The increase range can be a certain percentage of the current task number. For example, increase by 10% to 20%, or increase a fixed number of tasks, such as increasing 1 to 3.
[0103] 2. If any one or more index scores are lower than their baseline values, the system reduces the number of tasks. The reduction range can also be set as a certain percentage of the current task number. For example, reduce by 5% to 10%, or reduce 1 to 2 tasks.
[0104] 3. The setting of the baseline value should be adjusted according to the speaker's initial ability and training stage. It can be set relatively low initially and gradually increased as the speaker's ability improves to ensure a gradual increase in the training difficulty.
[0105] 4. To prevent the speaker from having difficulty adapting due to frequent adjustments, an adjustment period can be set. For example, evaluate once after each round of tasks is completed, or make adjustments after the average performance is stable in several consecutive tasks.
[0106] Through the multi-task parallel processing mechanism, the challenges of information processing and response for the speaker in a highly stressful environment are simulated, not limited to the Q&A session, but also covering various scenarios such as information integration, view elaboration, data interpretation, and instant feedback during the speech process.
[0107] This embodiment can generate multiple speech tasks based on the speech content, allocate multiple target tasks to the speaker according to the priority. After receiving the multiple target tasks, the speaker needs to respond to each target task to generate the corresponding information flow. By analyzing the information flow corresponding to each target task, the quick thinking and reaction ability of the speaker in the face of sudden situations and complex problems can be evaluated, so as to generate corresponding oratory expression suggestions. Based on the in-depth analysis of the speech content and the real-time information flow, multiple tasks are processed simultaneously. Through automated evaluation in multiple dimensions such as priority sorting, logical rigor, and emotional investment, the training intensity is dynamically adjusted to enhance the speaker's quick thinking agility, language organization, and information transmission efficiency in complex situations, and improve the comprehensive expressiveness of oratory skills. During this period, the speaker needs to process multiple target tasks in parallel, which increases the training difficulty of the speaker. The speaker is trained to maintain the clarity, coherence, and persuasiveness of language in multi-task processing, ensure the effective transmission of information, and improve the quality of oratory expression training.
[0108] In some embodiments, a speech analysis system is further provided, which executes the speech analysis method of multi-task parallel processing as in Embodiment 1. The functions of each module in the speech analysis system can be referred to the corresponding descriptions in the above method, and will not be elaborated here.
[0109] In some embodiments, an electronic device is further provided, such as Figure 2 the structural block diagram of the electronic device shown. The electronic device includes: a memory 100 and a processor 200. The memory 100 stores a computer program that can run on the processor 200. When the processor 200 executes the computer program, it implements the speech analysis method of multi-task parallel processing in the above embodiments. The number of the memory 100 and the processor 200 can be one or more.
[0110] The electronic device further includes:
[0111] a communication interface 300, which is used to communicate with external devices and perform data interaction and transmission. In addition, an API interface is provided to facilitate integration with other educational platforms or third-party applications, expanding the application scope.
[0112] If the memory 100, the processor 200, and the communication interface 300 are implemented independently, the memory 100, the processor 200, and the communication interface 300 can be interconnected via a bus to complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 2 only a thick line is used to represent it in Figure 2 , but it does not mean that there is only one bus or one type of bus.
[0113] Optionally, in a specific implementation, if the memory 100, the processor 200, and the communication interface 300 are integrated on a single chip, the memory 100, the processor 200, and the communication interface 300 can complete communication with each other through an internal interface.
[0114] In some embodiments, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the method provided in the embodiments of the present invention.
[0115] In some embodiments, a chip is also provided, which includes a processor for calling and running instructions stored in a memory, so that a communication device installed with the chip executes the method provided in the embodiments of the present invention.
[0116] In some embodiments, a chip is also provided, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected to each other through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiments of the invention.
[0117] It should be understood that the above-mentioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor that supports the advanced RISC machines (ARM) architecture.
[0118] Furthermore, optionally, the above-mentioned memory can include a read-only memory and a random access memory, and can also include a non-volatile random access memory. The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0119] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another.
[0120] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0121] Furthermore, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of these features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0122] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various changes or substitutions, and these should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A speech analysis method for multi-task parallel processing, characterized in that, including: Obtain the speech content and generate multiple speech tasks based on the speech content; among them, preprocess the speech content, including segmenting / sentencing the speech content through a natural language analysis algorithm, extracting keywords of each segment / sentence through a natural language analysis algorithm, analyzing the core purpose of each segment / sentence through an AI algorithm, and dividing multiple speech tasks according to the analyzed core purpose. The speech tasks include an introduction task, a current situation analysis task, a solution task, and an action task; Determine multiple target tasks to be assigned according to the priorities of all the speech tasks, and assign the multiple target tasks to the speaker simultaneously at a specified time point; among them, the method for determining the priorities is: Identify multiple information units from the speech content; count the number of the information units associated with each speech task; evaluate the information density included in each information unit to obtain the load of the information unit. The load of the information unit is quantified by information entropy or the number of key information points, reflecting the amount of information or importance carried by the information unit; calculate the information load corresponding to each speech task according to the number of the information units and the load of the information unit, and determine the priorities according to the information load; Obtain the information flow generated when the speaker processes the multiple target tasks in parallel. In the case of obtaining the information flow generated when the speaker processes the multiple target tasks in parallel, identify the key features of each information flow, find the corresponding target tasks of each information flow according to the key features, bind the information flow to the corresponding target task, and perform an oratory performance analysis on each information flow to generate corresponding expression improvement suggestions.
2. The speech analysis method for multi-task parallel processing according to claim 1, characterized in that, The oratory performance analysis on each information flow includes: Extract the index features associated with a preset index from the information flow; Calculate the index scores of each information flow on the preset index according to the index features.
3. The speech analysis method for multi-task parallel processing according to claim 2, wherein The preset index includes topic coherence, emotional engagement, information density, audience interactivity, and logic.
4. The speech analysis method for multi-task parallel processing according to claim 2, wherein The generation of the corresponding expression improvement suggestions includes: Obtain the speech target corresponding to each target task; Judge whether the index score corresponding to each target task reaches its corresponding speech target. In the case that the index score of any target task does not reach its corresponding speech target, generate the corresponding expression improvement suggestion according to the index score.
5. The speech analysis method for multi-task parallel processing according to claim 2, characterized in that It also includes: Dynamically adjust the number of target tasks assigned to the speaker according to the index score.
6. A speech analysis system, characterized in that, Execute the speech analysis method for multi-task parallel processing according to any one of claims 1 to 5.
7. An electronic device, characterized in that, including: A processor and a memory. Instructions are stored in the memory and loaded and executed by the processor to implement the speech analysis method for multi-task parallel processing according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, it implements the speech analysis method for multi-task parallel processing according to any one of claims 1 to 5.
Citation Information
Patent Citations
Oral talent training method based on document editing and communication
CN118135856A