Speech data processing method, system and device based on optical signal transmission and medium

By using optical signal transmission technology, speech data is encoded into optical signals for processing and decoding to generate suggestions for improving expression. This solves the problem of feedback delay in existing methods and enables real-time and personalized speech skill optimization.

CN118282511BActive Publication Date: 2026-03-24CHINA NEW LINE EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-03
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In existing methods for training oral expression and thinking skills, the analysis and feedback process takes a long time, and broadband network anomalies cause feedback delays, affecting the timeliness of feedback.

Method used

Using optical signal transmission technology, the speech data is converted into optical signals through an encoding model, and then processed and decoded to generate suggestions for improving the presentation, guiding speakers to optimize their presentation skills.

Benefits of technology

It improves feedback speed, enables instant and targeted presentation skills guidance, and solves the problem of delayed feedback in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118282511B_ABST
    Figure CN118282511B_ABST
Patent Text Reader

Abstract

The application provides a speech data processing method, system, device and medium based on optical signal transmission, which comprises the following steps: obtaining speech data, quantitatively analyzing the speech skill indexes of a speaker according to the speech data to obtain index quantitative values; encoding the index quantitative values into optical signals based on a preset encoding model for transmission; performing signal processing on the optical signals in the transmission process according to a preset speech target to obtain output signals; performing signal change analysis on the output signals to extract specified signals that have changed; and decoding the specified signals based on a preset decoding model to generate expression improvement suggestions to guide the speaker to adjust the speech skills. The application replaces the traditional network transmission with optical signal transmission, thereby improving the feedback speed. In addition, the quantitative analysis of the speech skill indexes is combined with multi-dimensional data processing generation, the signal processing strategy is dynamically adjusted, the speaker is guided in real time and in a targeted manner, and the training effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of public speaking training, and in particular to a method, system, device, and medium for processing speech data based on optical signal transmission. Background Technology

[0002] Current methods for improving public speaking and critical thinking skills typically require in-depth analysis of the entire speech, sometimes even necessitating real-time feedback of the analysis process and results to the speaker. However, due to the sheer volume of data analyzed and the use of broadband networks for transmission, this process inevitably takes time, leading to delays in feedback. Furthermore, if the broadband network malfunctions, data transmission speed is affected, making it impossible to provide the speaker with the necessary information promptly, thus impacting the timeliness of feedback. Summary of the Invention

[0003] This invention provides a method, system, device, and medium for processing speech data based on optical signal transmission to solve the problems existing in related technologies. The technical solution is as follows:

[0004] In a first aspect, embodiments of the present invention provide a speech data processing method based on optical signal transmission, comprising:

[0005] Acquire speech data, and conduct quantitative analysis on the speech ability indicators of the performers based on the speech data to obtain quantitative values ​​of the indicators;

[0006] The index quantization value is encoded into an optical signal for transmission based on a preset encoding model;

[0007] The optical signal during transmission is processed according to the preset speech target to obtain the output signal;

[0008] Perform signal change analysis on the output signal to extract the specified signal that has changed;

[0009] Based on a preset decoding model, the specified signal is decoded to generate suggestions for improvement, guiding speakers to adjust their presentation skills.

[0010] In one implementation, the performance indicators for public speaking include influence, persuasiveness, logical coherence, richness of emotional expression, and adaptability.

[0011] In one implementation, where the presentation ability indicator is an influence indicator, the quantitative analysis includes:

[0012] The number of standards used to assess influence is determined based on pre-defined influence assessment rules;

[0013] Feature extraction was performed on the speech data under various standards to obtain the key features corresponding to each standard;

[0014] Influence analysis is conducted based on key features to obtain influence scores for each standard;

[0015] The influence score corresponding to each standard is calculated, and the product of the weights corresponding to each standard is used to obtain the quantitative value of the influence index.

[0016] In one implementation, when the presentation ability indicator is logical coherence, the quantitative analysis includes:

[0017] The speech data is segmented into multiple logical units;

[0018] Perform coherence analysis on each logical unit to obtain a discontinuity score;

[0019] Based on the number of logical units in the speech data and the incoherence score corresponding to each logical unit, a quantitative value of logical coherence is calculated.

[0020] In one implementation, when the presentation ability indicator is adaptive, the quantitative analysis includes:

[0021] Determine the scenarios and number of scenarios to be used for assessing adaptability based on the pre-set training needs and objectives;

[0022] Adaptive analysis was performed on the speech data in various contexts to obtain adaptive scores for each context;

[0023] The product of the adaptability score for each scenario and the corresponding weight for each scenario is used to obtain the quantitative value of the adaptability index.

[0024] In one implementation, signal processing includes one or a combination of two or more of the following: separation, extraction, amplification, superposition, and filtering.

[0025] In one implementation, signal change analysis of the output signal includes:

[0026] The output signal is compared with the optical signal to determine whether the output signal has changed. Signal changes include signal loss and signal distortion.

[0027] When the output signal changes, the changed signal is extracted and marked as a specified signal.

[0028] Secondly, embodiments of the present invention provide a speech data processing system that performs the speech data processing method based on optical signal transmission as described above.

[0029] Thirdly, embodiments of the present invention provide an electronic device comprising a memory and a processor. The memory and the processor communicate with each other via an internal connection path. The memory stores instructions, and the processor executes the instructions stored in the memory. When the processor executes the instructions stored in the memory, it causes the processor to perform the method described in any of the above embodiments.

[0030] Fourthly, embodiments of the present invention provide a computer-readable storage medium that stores a computer program, wherein when the computer program is run on a computer, the methods in any of the embodiments described above are executed.

[0031] The advantages or beneficial effects of the above technical solutions include at least the following:

[0032] This invention encodes the speaker's presentation data into optical signals for transmission. During transmission, the optical signals are processed and decoded according to preset presentation goals. Based on the decoded data, corresponding suggestions for improvement are generated to guide the speaker in optimizing their presentation skills. The use of optical signal transmission instead of traditional network transmission increases transmission speed, and the optimization of the speaker's techniques during transmission improves feedback speed and solves the real-time problem. Furthermore, the quantitative analysis of presentation ability indicators combines multi-dimensional data processing and personalized suggestion generation. By dynamically adjusting signal processing strategies, it achieves immediate and targeted guidance for the speaker, a capability not found in traditional analysis methods.

[0033] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the invention will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0034] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in the invention and should not be construed as limiting the scope of the invention.

[0035] Figure 1 This is a flowchart illustrating the speech data processing method based on optical signal transmission according to the present invention.

[0036] Figure 2 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0037] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0038] Example 1

[0039] This embodiment provides a speech data processing method based on optical signal transmission, such as... Figure 1 As shown, the method includes:

[0040] Step S1: Obtain speech data, and perform quantitative analysis on the speech ability indicators of the performers based on the speech data to obtain quantitative values ​​of the indicators;

[0041] Step S2: Encode the index quantization value into an optical signal for transmission based on a preset encoding model;

[0042] Step S3: Perform signal processing on the optical signal during transmission according to the preset speech target to obtain the output signal;

[0043] Step S4: Perform signal change analysis on the output signal and extract the specified signals that have changed;

[0044] Step S5: Decode the specified signal based on the preset decoding model and generate suggestions for improving the presentation to guide the speaker in adjusting their presentation skills.

[0045] The speech data primarily includes the speaker's verbal and nonverbal information. Verbal information can be obtained through speech recognition, while nonverbal information includes facial expressions, body movements, eye movement data, and electroencephalogram (EEG) data. Speech data can be real-time data generated during the speaker's current speech, historical data, or data reflecting audience feedback through facial expressions, body language, or verbal communication during the speech.

[0046] A speaker's speaking ability can be demonstrated from different dimensions during the presentation. Therefore, by pre-establishing multiple indicators of speaking ability and analyzing the speaker's scores on each indicator, a comprehensive analysis of the speaker's presentation process can be conducted from multiple dimensions, resulting in a more comprehensive analysis of the speaker's speaking ability and the provision of more comprehensive guidance and suggestions.

[0047] In this embodiment, the speech ability indicators include influence indicators, persuasiveness indicators, logical coherence, richness of emotional expression, and adaptability indicators. Influence indicators primarily evaluate the speaker's influence; persuasiveness indicators primarily evaluate the persuasiveness of the content delivered; logical coherence primarily evaluates the speaker's logical flow and fluency; richness of emotional expression primarily evaluates the speaker's ability to convey emotions across different dimensions; and adaptability indicators primarily evaluate the speaker's ability to adjust and adapt their speech style and content in different situations.

[0048] When public speaking ability is used as an indicator of influence, the process of conducting influence analysis on public speaking data and quantifying the results to obtain corresponding quantitative values ​​includes:

[0049] Step S11: Determine the number of standards used to assess influence based on preset influence assessment rules;

[0050] Step S12: Extract features from the speech data under each standard to obtain the key features corresponding to each standard;

[0051] Step S13: Conduct influence analysis based on key features to obtain the influence score corresponding to each standard;

[0052] Step S14: Calculate the product of the influence score corresponding to each standard and the weight corresponding to each standard to obtain the quantitative value of the influence index.

[0053] The quantitative analysis of influence indicators includes determining the number of evaluation criteria, feature extraction, scoring, and weight calculation. Specifically, the evaluation criteria are first determined according to the preset influence evaluation rules. Then, features are extracted from the speech data, such as the authority of the language and the intonation of the voice. Scores are then assigned based on these features, and the quantitative value of the influence indicator is calculated by combining the weights of the criteria.

[0054] Influence assessment rules can be pre-set according to actual circumstances, and multiple standards for assessing influence can be determined based on different rules. Therefore, influence assessment rules not only define the content of the standards for assessing influence, but also the number of standards defined. Essentially, a speaker can establish multiple standards for assessing influence based on actual needs.

[0055] Assuming the rules for assessing influence are defined by experts, the standards used to evaluate influence are a comprehensive and relevant set of assessment criteria defined by experts in the field based on their experience and expertise. Influence analysis of the speech data is then performed according to these standards; that is, experts in the field assess the speaker's level of influence based on the speech data and manually score the speaker's influence to obtain an influence score under these standards.

[0056] Since the impact analysis process differs under different standards, it is necessary to conduct impact analysis on the speech data separately under each standard to obtain the corresponding impact score for each standard.

[0057] If the rules for assessing influence are based on historical data, then the standards used to assess influence are those extracted from a large amount of historical data. For example, common key characteristics that give a speaker influence can be extracted from a large amount of historical data. These common key characteristics can be body language, facial features, or language features, or a combination of multiple characteristics can be used to reflect the degree of influence of the speaker's speech. These common key factors are then used as the assessment standards.

[0058] Under this standard, features are extracted from the speaker's speech data to determine whether the speaker possesses common key features. Since different common key features result in different levels of influence, the speaker's level of influence is scored using a specified scoring function. For example, a corresponding weight is set for each common key feature, and the speaker's influence score is obtained by multiplying all common key features by their corresponding weights and summing them up.

[0059] Assuming the criteria used to assess influence are expert analysis and historical data, the number of criteria is 2. A weight is assigned to each criterion, and the influence score for each criterion is calculated, along with the sum of the products of their respective weights, to obtain the quantitative value of the influence index. The quantitative value of the influence index is calculated using the following formula:

[0060] ;in,

[0061] Q is the number of criteria used to assess influence, ω q It is the weight of the q-th standard, IMP q (T) q The speaker's influence score on the q-th criterion is evaluated using a specified scoring function. The scoring function can be a logistic function (such as the sigmoid function) to represent the saturation effect of influence on performance; the principle of this function is already publicly available and will not be described in detail here. The influence score can also be obtained through manual scoring.

[0062] In the context of persuasiveness indicators, speech ability is used to analyze the persuasiveness of speech data. The results are then quantified to obtain the corresponding persuasiveness index value. The calculation process is the same as that for influence indicators. Similarly, standards for evaluating explanatory power are pre-established. The speaker's persuasiveness score is then assessed under each standard. The persuasiveness index value is calculated by summing the products of the explanatory power scores under each standard and the corresponding weights for each standard. The formula for calculating the persuasiveness index value is as follows:

[0063] ;in,

[0064] Q is the number of criteria used to evaluate persuasiveness, θ r It is the weight of the r-th criterion, CONV r (E) r The persuasiveness score of the speaker on the r-th criterion is evaluated using a specified scoring function. The scoring function can be a logistic function, the principle of which is already disclosed in existing technology and will not be described in detail here. Alternatively, the persuasiveness score can also be obtained through human scoring.

[0065] When the presentation ability indicator is logical coherence, the quantitative analysis includes:

[0066] Step S15: Perform text segmentation on the speech data to obtain multiple logical units;

[0067] Step S16: Perform coherence analysis on each logical unit to obtain a discontinuity score;

[0068] Step S17: Calculate the logical coherence quantification value based on the number of logical units in the presentation data and the incoherence score corresponding to each logical unit.

[0069] The specific steps of logical coherence quantification analysis include text segmentation into logical units, scoring the coherence of each logical unit, and calculating a quantification value based on the degree of incoherence. In practice, logical units are determined through methods such as semantic analysis and syntactic analysis, the coherence of each unit is analyzed, and finally, the incoherence scores of all logical units are summarized to obtain the quantification value.

[0070] Identifying the logical units within a speech is a crucial step, as they form the basis for assessing logical coherence. A logical unit typically refers to a paragraph in a speech that expresses a complete idea or argument.

[0071] The speech data is segmented into multiple logical units using methods including semantic segmentation, syntactic analysis, content structure tagging, and logical relation identification. Semantic segmentation utilizes natural language processing techniques to divide the speech content into independent logical units through semantic analysis. Syntactic analysis primarily identifies the main and subordinate clause relationships in the speech, considering the preceding main clause as the starting point of the logical unit, and the subordinate clause as a supplement or explanation of the main clause. Content structure tagging identifies structural markers in the speech, such as "firstly," "secondly," "however," and "therefore," which typically indicate the beginning of a new logical unit. Logical relation identification techniques determine the boundaries of logical units by analyzing the logical relationships in the speech content, such as causality, contrast, and parallelism; the points of change in each logical relationship can serve as the dividing lines of logical units.

[0072] In addition, audience feedback can be considered, especially the parts of the speech that elicit a response from the audience; these points may also be the boundaries of logical units.

[0073] After defining the boundaries of the logical units, a coherence analysis is performed on each unit. AI technology is used to evaluate the coherence of the context within each unit, and a corresponding score is given for the degree of incoherence. The incoherence score can also be obtained through manual evaluation.

[0074] Based on the number of logical units in the presentation data and the incoherence score corresponding to each logical unit, a quantitative value for logical coherence is calculated, and the quantitative formula is as follows:

[0075] ;

[0076] Where N is the number of logical units in the presentation content, and D i It is the discontinuity score of the i-th logical unit.

[0077] When the metric for presentation ability is the richness of emotional expression, its quantification formula is:

[0078] ;

[0079] In the formula, M is the number of dimensions used to assess emotional expression, and φ j It is the weight of the j-th dimension, EMO 2 j It is the emotional expression intensity score of the j-th dimension. The emotional expression intensity score can be obtained through manual scoring or through AI technology analysis.

[0080] The content and number of dimensions used to assess emotional expression can be set through preset rules. The dimensions used to assess emotional expression include theoretical dimensions, expert consultation dimensions, and data analysis dimensions.

[0081] The theoretical dimension is based on theories from fields such as psychology, communication studies, and performing arts to identify the key elements of emotional expression; the intensity score of emotional expression in this dimension is determined by analyzing whether the speech data contains these key elements.

[0082] The expert consultation dimension involves consulting experts in the fields of emotional expression and public speaking training to understand the emotional expression elements they consider important; and determining the emotional expression intensity score in this dimension by analyzing whether the speech data contains these emotional expression elements.

[0083] The data analysis dimension identifies key factors that influence the effectiveness of emotional expression by analyzing historical speech data and audience feedback; it determines the intensity score of emotional expression in this dimension by analyzing whether the speech data contains these key elements.

[0084] In the context of public speaking or oratory training, adaptability refers to a speaker's ability to flexibly adjust their speaking style, content, tone, and expression to suit different audiences, occasions, cultural backgrounds, or specific situations. A highly adaptable speaker can quickly adjust to factors such as the atmosphere, audience reactions, and technical glitches, maintaining the coherence, appeal, and effectiveness of their speech. Adaptability reflects a speaker's sensitivity to environmental changes and their ability to respond quickly, and is one of the important indicators for measuring their overall competence.

[0085] Different speaking situations have unique characteristics and requirements, so the adaptability standards will also change accordingly:

[0086] In situations involving diverse audiences (such as professionals, the general public, and teenagers), different audience groups have varying knowledge backgrounds, interests, and expectations, leading to different requirements for the content and style of the speech. In situations involving the nature of the occasion (formal meetings, academic lectures, public speaking, online live streams, etc.), different requirements exist regarding the form, etiquette, and depth of the speech. In situations involving cultural factors (audiences from different cultural backgrounds have varying levels of acceptance of speaking styles, humor, and direct or indirect expression), the length of the speech, the physical conditions of the venue (such as lighting and sound), and emergency situations all influence the presentation of the speech. In situations involving the purpose of the speech (education, persuasion, motivation, entertainment, etc.), different speech objectives require speakers to adopt different strategies in content arrangement, emotional engagement, and information delivery.

[0087] Therefore, the setting of adaptability indicators must take these specific situational factors into account to ensure that the evaluation criteria can accurately reflect the speaker's ability to adjust the speaking strategy to achieve the best communication effect in a specific situation.

[0088] When the presentation ability indicator is adaptive, the quantitative analysis includes:

[0089] Step S18: Determine the scenarios and number of scenarios to be used for assessing adaptability based on the preset training needs and objectives;

[0090] Step S19: Perform adaptive analysis on the speech data in each context to obtain an adaptive score for each context;

[0091] Step S20: Calculate the product of the adaptability score for each scenario and the weight corresponding to each scenario to obtain the quantitative value of the adaptability index.

[0092] The quantifiable value of the adaptability index can be evaluated using the following formula:

[0093] ;

[0094] Where K is the number of scenarios used to assess adaptability, ξk It is the weight of the k-th scenario, ADAPT k (S) k () refers to the assessment of a speaker's adaptability score in the k-th situation. The adaptability score can be calculated by a function, which can be a logistic function or by manually scoring the adaptability in different situations.

[0095] The detailed process of adaptive index quantification analysis includes determining the evaluation context, conducting adaptive analysis under different contexts, calculating the adaptive score and weighted product. First, the evaluation context is determined based on training needs and objectives. Then, the adaptability of the speech data under each context is scored. Finally, the quantitative value of the adaptive index is calculated based on the context weights.

[0096] In this embodiment, after extracting the quantized values ​​of each indicator, the quantized values ​​are encoded and converted according to a preset encoding scheme. During the encoding and conversion process, in addition to converting the quantized values ​​of each indicator, the speech content or speech features associated with the quantized values ​​are also converted. The speech data is converted into optical signals through a preset encoding model. The encoding model employs advanced modulation techniques, such as Orthogonal Frequency Division Multiplexing (OFDM), to map the quantized values ​​to optical signal parameters such as light intensity, frequency, or phase, achieving efficient data encoding. The encoding model design follows the principle of minimum distortion, ensuring high information density and ease of decoding.

[0097] The key to the encoding model's encoding rules lies in transforming lengthy speech data into a concise expression. The encoded optical signals represent various dimensions and key characteristics of the speaker's eloquence. These optical signals can be transmitted via optical fiber, significantly increasing data transmission speed and reducing delays caused by broadband network lag. Furthermore, signal processing can be applied to the optical signals to adjust the speaker's delivery techniques and information transmission effectiveness.

[0098] The encoded optical signal undergoes corresponding signal processing, including one or more methods such as separation, extraction, amplification, superposition, and filtering, or a combination of two or more methods. This embodiment pre-establishes a mapping relationship between different presentation goals and different signal processing methods; this mapping relationship can be determined in advance through extensive experimentation. This is equivalent to applying the concept of signal change in an optical path as it passes through different optical elements to the presentation adjustment method in this embodiment, enabling the presenter to adjust their presentation skills through optical signal processing.

[0099] After signal processing, the optical signal is converted into an output signal. The output signal is then compared with the input optical signal to determine if the output signal has changed. Changes include signal loss and signal distortion. The changed signal is then marked as a designated signal.

[0100] The system decodes a specified signal based on a pre-defined decoding model, providing suggestions for improvement. The decoding model utilizes matching demodulation techniques combined with advanced signal processing algorithms, such as digital signal processing (DSP), to recover the original quantized values ​​from the received optical signal, ensuring data accuracy. The decoding model pre-defines rules linking signal variations to speech skill improvements; for example, signal distortion suggests improved speech clarity, while signal loss suggests reducing redundant content. The model automatically matches appropriate suggestions for improvement based on the type of signal variation. The decoding model design emphasizes robustness and accuracy, ensuring correct decoding even under weak or interfering signal conditions.

[0101] This embodiment pre-binds the correlation between signal changes and speaker improvement points through a decoding model. When the output signal changes, the specified signal that has changed is decoded, and corresponding improvement suggestions are generated based on the signal changes of the specified signal. This leads to a feedback report that points out the speaker's strengths and areas for improvement in their presentation. The correlation between signal distortion and presentation improvement can be as follows:

[0102] If some signals are distorted during processing, it may mean that some of the speaker's expressions are not clear or ambiguous. In this case, the speaker can be guided to improve in this regard, such as optimizing language choices and sentence structure.

[0103] If some signals are lost during processing, it indicates that the speaker's expression is relatively redundant. In this case, redundant expressions can be identified and eliminated, thereby improving the efficiency of the speech and the audience's focus.

[0104] If certain information remains unchanged or is enhanced during signal processing, this indicates that these are effective parts of the speech. At this point, the speaker can be reminded to identify these strengths and continue to maintain them in future speeches.

[0105] At the same time, suggestions for improvement can be mapped to the corresponding paragraphs of the original speech content in the speech data to generate new speech content to guide speakers in adjusting their speech delivery.

[0106] In some embodiments, the speaker's speech data can be quantitatively analyzed using natural language processing and machine learning algorithms to generate corresponding feedback suggestions. These suggestions include training recommendations for improvement needed by the speaker. These training suggestions are used as a data source for encoding model optimization, signal processing optimization, and decoding optimization, making the analysis of the encoding model, signal processing mapping relationship, and decoding model more accurate, and resulting in a more accurate final feedback report. The methods for analyzing the speaker's speech content using natural language processing and machine learning algorithms to obtain feedback suggestions are already disclosed in the prior art and will not be described in detail here.

[0107] Optical signal transmission converts quantized values ​​into optical signals using an encoding model and then transmits them via transmission media such as optical fibers. Signal processing includes operations such as separation, extraction, amplification, superposition, and filtering, optimizing the signal according to the transmission objectives. Specific technical implementations involve optical modulation and demodulation techniques, signal processing devices, and algorithms.

[0108] This embodiment encodes the speaker's presentation data into optical signals for transmission. During transmission, the optical signals are processed and decoded according to preset presentation goals. Based on the decoded data, corresponding suggestions for improvement are generated to guide the speaker in optimizing their presentation skills. The use of optical signal transmission instead of traditional network transmission increases transmission speed, and the optimization of the speaker's presentation techniques during transmission improves feedback speed.

[0109] Example 2

[0110] This embodiment provides a speech data processing system that executes the speech data processing method based on optical signal transmission as described in Embodiment 1. The system design is compatible with common educational technology standards, such as SCORM or LTI, ensuring seamless integration with existing teaching platforms. In terms of quantitative indicators, natural language processing (NLP) and machine learning algorithms are used to train a model using historical speech data to accurately assess the speaker's level in terms of influence, logical coherence, etc., and the effectiveness is demonstrated through actual teaching cases.

[0111] The functions of the system in this embodiment of the invention can be found in the corresponding descriptions in the above methods, and will not be repeated here.

[0112] In some embodiments, an electronic device, such as Figure 2 The illustrated block diagram of the electronic device includes a memory 100 and a processor 200. The memory 100 stores a computer program that can run on the processor 200. When the processor 200 executes the computer program, it implements the speech data processing method based on optical signal transmission described in the above embodiments. The number of memories 100 and processors 200 can be one or more.

[0113] The electronic device also includes:

[0114] The communication interface 300 is used to communicate with external devices and perform data exchange and transmission.

[0115] If the memory 100, processor 200, and communication interface 300 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 2 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0116] Optionally, in a specific implementation, if the memory 100, processor 200, and communication interface 300 are integrated on a single chip, then the memory 100, processor 200, and communication interface 300 can communicate with each other through an internal interface.

[0117] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method provided in this invention.

[0118] This invention also provides a chip, which includes a processor for calling and executing instructions stored in a memory, causing a communication device on which the chip is installed to perform the method provided in this invention.

[0119] This invention also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in this invention.

[0120] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting the Advanced Reduced Instruction Set Computing (RISC) machine (ARM) architecture.

[0121] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0122] In the above embodiments, implementation can be achieved, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0123] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0124] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0125] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in the present invention, and these should all be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A speech data processing method based on optical signal transmission, characterized in that, include: Acquire speech data, and conduct quantitative analysis of the speaker's speech ability indicators based on the speech data to obtain quantitative values ​​of the indicators; The quantized values ​​of the indicators are encoded into optical signals for transmission based on a preset encoding model; The encoding conversion process also includes the conversion of speech content or speech characteristics associated with the quantitative values ​​of the indicators; The optical signal is processed according to the preset speech target to obtain the output signal; Perform signal change analysis on the output signal to extract the specified signal that has changed; The specified signal is decoded based on a preset decoding model to generate suggestions for improving the presentation skills of the speaker. The quantified values ​​of the indicators are mapped to optical signal parameters, establishing a mapping relationship between different speech targets and different signal processing methods; The decoding model pre-binds the correlation between signal changes and speaker improvement points. When the output signal changes, the specified signal that has changed is decoded, and corresponding suggestions for improvement are generated based on the signal changes of the specified signal. This leads to a corresponding feedback report that points out the speaker's strengths and areas for improvement in their presentation.

2. The speech data processing method based on optical signal transmission according to claim 1, characterized in that, The speech performance indicators include influence, persuasiveness, logical coherence, richness of emotional expression, and adaptability.

3. The speech data processing method based on optical signal transmission according to claim 2, characterized in that, When the presentation ability indicator is the influence indicator, the quantitative analysis includes: The number of standards used to assess influence is determined based on pre-defined influence assessment rules; Feature extraction is performed on the speech data under each standard to obtain the key features corresponding to each standard; Influence analysis is performed based on the aforementioned key features to obtain the influence scores corresponding to each standard; The influence score corresponding to each standard is calculated, and the product of the weights corresponding to each standard is used to obtain the quantitative value of the influence index.

4. The speech data processing method based on optical signal transmission according to claim 2, characterized in that, When the presentation ability indicator is logical coherence, the quantitative analysis includes: The speech data is segmented into multiple logical units; A coherence analysis is performed on each of the aforementioned logical units to obtain a discontinuity score; Based on the number of logical units in the speech data and the incoherence score corresponding to each logical unit, a quantitative value of logical coherence is calculated.

5. The speech data processing method based on optical signal transmission according to claim 2, characterized in that, When the presentation ability indicator is described as adaptability, the quantitative analysis includes: Determine the scenarios and number of scenarios to be used for assessing adaptability based on the pre-set training needs and objectives; Adaptive analysis was performed on the speech data under various scenarios to obtain an adaptation score for each scenario; The product of the adaptability score for each scenario and the corresponding weight for each scenario is used to obtain the quantitative value of the adaptability index.

6. The speech data processing method based on optical signal transmission according to claim 1, characterized in that, The signal processing includes one or a combination of two or more of the following: separation, extraction, amplification, superposition, and filtering.

7. A speech data processing system, characterized in that, Perform the speech data processing method based on optical signal transmission as described in any one of claims 1 to 6.

8. An electronic device, characterized in that, include: A processor and a memory, wherein the memory stores instructions that are loaded and executed by the processor to implement the speech data processing method based on optical signal transmission as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the speech data processing method based on optical signal transmission as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • High speed optical transmission system and method based on TCM-64QAM code modulation

    CN102088317A

  • Intelligent speech training system and method

    CN112232127A