Method and device for pre-detecting listening test content

By voice recognition and comparison of the front part, boundary label and post part of the listening test content, the examination problems caused by CD file errors are solved to ensure that the examination is carried out normally.

CN120340495APending Publication Date: 2025-07-18刘江华
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510588365.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the language listening test, due to the erroneous format, defects or inconsistent content of the CD or USB disk file, the examination audio files cannot be played normally, which may cause the examination accident.

Method used

By picking up the front part, the first dividing label and/or the back part of the listening test content for speech recognition, the speech recognition results are generated, and compared with the pre-acquisitioned text content, the correctness of the listening test content and whether there is any lag.

Benefits of technology

Without contacting the physical part of the listening test content, ensure that the test content is correct and check whether there is any playback lag to avoid the occurrence of test accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340495A_ABST
    Figure CN120340495A_ABST
Patent Text Reader

Abstract

The invention discloses a hearing test content pre-detection method and device. The method comprises the following steps: picking up a first boundary label and / or a rear part in a front part and a hearing entity part of the hearing test content; performing voice recognition on the front part, the first boundary label and / or the rear part to generate a voice recognition result; detecting whether the listening test content is correct or not according to the voice recognition result, the pre-acquired literal content of the front part, the pre-acquired literal content of the first boundary label and / or the pre-acquired literal content of the rear part; and if the listening test content is correct, detecting whether the listening test content is lagged or not according to the voice recognition result. According to the invention, on the premise that actual examination questions of the listening examination content are not contacted, problems of listening examination audio files or listening examination content errors in the prior art can be completely eradicated, and normal listening examination can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of sound reproduction technology, and particularly relates to a method and device for pre-detecting the content of a listening exam. Background Art

[0002] In the prior art, the broadcast content of a language listening exam must be distributed to each test site through carriers such as optical discs (or USB flash drives). For security reasons, no one can hear the content before the official exam starts. In this case, there may be exam accidents due to errors in the listening content or problems with the listening exam audio file. The above-mentioned errors in the listening content specifically include:

[0003] 1. The file format of the optical disc (USB flash drive) is incorrect;

[0004] 2. The optical disc has defects, resulting in playback stuttering or being unable to play at all;

[0005] 3. The listening content does not match the exam: (recording errors during optical disc (USB flash drive) burning, copying errors during copying, or incorrect distribution of the optical disc (USB flash drive)). Summary of the Invention

[0006] An object of the present invention is to provide a method for pre-detecting the content of a listening exam, which can detect the content of the listening exam before the exam without exposing the listening entity part (the real exam questions) of the listening exam content, so as to eliminate the possibility of exam accidents caused by errors in the listening content or problems with the listening exam audio file in the prior art and ensure the normal progress of the listening exam.

[0007] Another object of the present invention is to provide a device for pre-detecting the content of a listening exam. Still another object of the present invention is to provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned method for pre-detecting the content of a listening exam are implemented. Still another object of the present invention is to provide a readable medium, on which a computer program is stored, and when the computer program is executed by the processor, the steps of the above-mentioned method for pre-detecting the content of a listening exam are implemented.

[0008] To solve the technical problems in the background art of this application, the present invention provides the following technical solutions:

[0009] In a first aspect, the present invention provides a method for pre-detecting the content of a listening exam, the method including:

[0010] Pick up the preamble part, the first demarcation label in the listening entity part, and / or the postscript part of the listening exam content; wherein, the preamble part is the representation information of the listening exam content, the postscript part is the ending information of the listening exam content, and the first demarcation label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the exam questions in the listening entity part;

[0011] Perform speech recognition on the preamble part, the first demarcation label, and / or the postscript part to generate a speech recognition result;

[0012] Detect whether the listening exam content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first demarcation label, and / or the text content of the postscript part;

[0013] If the listening exam content is correct, detect whether there is any lag in the listening exam content according to the speech recognition result.

[0014] In some embodiments of the present invention, a pre-detection method for listening exam content further includes:

[0015] Set at least one second demarcation label in the listening exam content; the second demarcation label is used to divide the preamble part from the listening entity part of the listening exam content, and to divide the listening entity part from the postscript part;

[0016] Encrypt the listening entity part;

[0017] The picking up the preamble part and / or the postscript part of the listening exam content includes:

[0018] Determine the position of the second demarcation label;

[0019] Pick up the preamble part and / or the postscript part according to the position of the second demarcation label.

[0020] In some embodiments of the present invention, the detecting whether the listening exam content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first demarcation label, and / or the text content of the postscript part includes:

[0021] Compare the speech recognition result with the text content;

[0022] If the speech recognition result is consistent with the text content, determine that the listening exam content is correct;

[0023] Otherwise, determine that the listening exam content is incorrect.

[0024] In some embodiments of the present invention, the text content includes: the first demarcation label, the second demarcation label, the year, month, and name of the listening exam.

[0025] In some embodiments of the present invention, detecting whether there is a lag in the listening exam content according to the speech recognition result includes:

[0026] Detecting whether there is a lag in the listening exam content according to the first time interval between two adjacent sentences in the speech recognition result and a preset first threshold.

[0027] In some embodiments of the present invention, detecting whether there is a lag in the listening exam content according to the speech recognition result includes:

[0028] Determining at least one keyword in the speech recognition result;

[0029] Detecting whether there is a lag in the listening exam content by determining the second time interval between the keyword and the initial time of the preamble part and a preset second threshold.

[0030] In some embodiments of the present invention, detecting whether there is a lag in the listening exam content according to the speech recognition result further includes:

[0031] Determining at least two keywords in the speech recognition result;

[0032] Detecting whether there is a lag in the listening exam content by determining the third time interval between the two keywords and a preset third threshold.

[0033] In a second aspect, the present invention provides a pre-detection device for listening exam content, and the device includes:

[0034] A pickup module, configured to pick up the preamble part, the first demarcation label in the listening entity part, and / or the postscript part of the listening exam content; wherein, the preamble part is the characterization information of the listening exam content, the postscript part is the end information of the listening exam content, and the first demarcation label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the exam questions in the listening entity part;

[0035] A speech recognition result generation module, configured to perform speech recognition on the preamble part, the first demarcation label, and / or the postscript part to generate a speech recognition result;

[0036] A listening content correctness detection module, configured to detect whether the listening exam content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first demarcation label, and / or the text content of the postscript part.

[0037] A listening content stuttering detection module, which is used to detect whether there is stuttering in the listening exam content according to the speech recognition result when the listening exam content is correct.

[0038] In some embodiments of the present invention, a pre-detection device for listening exam content further includes:

[0039] A demarcation label setting module, which is used to set at least one second demarcation label in the listening exam content; the second demarcation label is used to divide the preamble part and the listening entity part of the listening exam content, and divide the listening entity part and the postscript part;

[0040] A listening entity part encryption module, which is used to encrypt the listening entity part;

[0041] The picking module includes:

[0042] A demarcation label position determination unit, which is used to determine the position of the second demarcation label;

[0043] A preamble part picking unit, which is used to pick the preamble part and / or the postscript part according to the position of the second demarcation label.

[0044] In some embodiments of the present invention, the listening content correctness detection module includes:

[0045] A comparison unit, which is used to compare the speech recognition result with the text content;

[0046] An exam content correct judgment unit, which is used to judge that the listening exam content is correct when the speech recognition result is consistent with the text content;

[0047] An exam content error judgment unit, which is used to otherwise judge that the listening exam content is wrong.

[0048] In some embodiments of the present invention, the text content includes: the first demarcation label, the second demarcation label, the year, month and name of the listening exam.

[0049] In some embodiments of the present invention, the listening content stuttering detection module includes:

[0050] A first listening content stuttering detection unit, which is used to detect whether there is stuttering in the listening exam content according to the first time interval between two adjacent sentences in the speech recognition result and a preset first threshold.

[0051] In some embodiments of the present invention, the listening content stuttering detection module includes:

[0052] A keyword determination first unit for determining at least one keyword in the speech recognition result;

[0053] A listening content lag detection second unit for determining a second time interval between the keyword and the initial time of the previous part and detecting whether there is a lag in the listening test content according to a preset second threshold.

[0054] In some embodiments of the present invention, the listening content lag detection module further includes:

[0055] A keyword determination second unit for determining at least two keywords in the speech recognition result;

[0056] A listening content lag detection third unit for determining a third time interval between the two keywords and detecting whether there is a lag in the listening test content according to a preset third threshold.

[0057] In a third aspect, the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of a method for pre-detecting listening test content are implemented.

[0058] In a fourth aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of a method for pre-detecting listening test content are implemented.

[0059] In a fifth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of a method for pre-detecting listening test content are implemented.

[0060] As can be seen from the above description, the embodiments of the present invention provide a method and device for pre-detecting listening test content. The corresponding method for pre-detecting listening test content includes: First, pick up the previous part, the first delimiter label in the listening entity part and / or the subsequent part of the listening test content; wherein, the previous part is the characterization information of the listening test content, the subsequent part is the end information of the listening test content, and the first delimiter label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the test questions in the listening entity part; perform speech recognition on the previous part, the first delimiter label and / or the subsequent part to generate a speech recognition result; then, detect whether the listening test content is correct according to the speech recognition result, the pre-obtained text content of the previous part and / or the text content of the subsequent part; if the listening test content is correct, detect whether there is a lag in the listening test content according to the speech recognition result.

[0061] In summary, the pre-detection method for the content of a listening exam provided by the present invention can detect the content of the listening exam before the exam without contacting the physical part of the listening exam content (the actual exam questions), so as to eliminate the possibility of exam accidents caused by errors in the listening content or problems with the listening exam audio file in the prior art, and ensure the normal progress of the listening exam. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 Flow diagram of a pre-detection method for the content of a listening exam in an embodiment of the present invention Figure 1 ;

[0064] Figure 2 Flow diagram of a pre-detection method for the content of a listening exam in an embodiment of the present invention Figure 2 ;

[0065] Figure 3 It is a flow diagram of step 100 of a pre-detection method for the content of a listening exam in an embodiment of the present invention;

[0066] Figure 4 It is a flow diagram of step 200 of a pre-detection method for the content of a listening exam in an embodiment of the present invention;

[0067] Figure 5 It is a flow diagram of step 300 of a pre-detection method for the content of a listening exam in an embodiment of the present invention Figure 1 ;

[0068] Figure 6 It is a flow diagram of step 300 of a pre-detection method for the content of a listening exam in an embodiment of the present invention Figure 2 ;

[0069] Figure 7 It is a schematic diagram of the working principle of a pre-detection system for the content of a listening exam in a specific embodiment of the present invention;

[0070] Figure 8 It is a block diagram of a pre-detection device for the content of a listening exam in an embodiment of the present invention Figure 1 ;

[0071] Figure 9Block diagram of a pre - detection device for the content of a listening exam in an embodiment of the present invention Figure 2 ;

[0072] Figure 10 Block diagram of the pickup module 10 in an embodiment of the present invention;

[0073] Figure 11 Block diagram of the listening content correctness detection module 30 in an embodiment of the present invention;

[0074] Figure 12 Block diagram of the listening content stutter detection module 40 in an embodiment of the present invention Figure 1 ;

[0075] Figure 13 Block diagram of the listening content stutter detection module 40 in an embodiment of the present invention Figure 2 ;

[0076] Figure 14 Schematic structural diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0077] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0078] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk storage, CD - ROM, optical storage, etc.) containing computer - usable program code.

[0079] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above - mentioned drawings are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. Without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0080] The information collected in the technical solution of this application is information and data authorized by the user or fully authorized by all parties, and the processing of related data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with the relevant laws, regulations, and standards of relevant countries and regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.

[0081] Provide corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, then enter the expert decision-making process.

[0082] The acquisition, storage, use, processing, etc. of data in the technical solution of this application comply with the relevant provisions of laws and regulations. Specifically:

[0083] First, the information collected is information and data authorized by the user or fully authorized by all parties, and the processing of related data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with the relevant laws, regulations, and standards of relevant countries and regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.

[0084] Second, provide corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, then enter the expert decision-making process.

[0085] An embodiment of the present invention provides a specific implementation manner of a method for pre-detecting the content of a listening test. Refer to Figure 1 , and this method specifically includes the following contents:

[0086] Step 100: Pick up the front part, the first demarcation label in the listening entity part, and / or the rear part of the listening test content; wherein, the front part is the characterization information of the listening test content, the rear part is the end information of the listening test content, and the first demarcation label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the question content in the listening entity part;

[0087] Step 200: Perform speech recognition on the front part, the first demarcation label, and / or the rear part to generate a speech recognition result;

[0088] Step 300: Detect whether the listening test content is correct according to the speech recognition result, the pre-acquired text content of the front part, the text content of the first demarcation label, and / or the text content of the rear part;

[0089] Step 400: When the listening test content is correct, detect whether there is any lag in the listening test content according to the speech recognition result.

[0090] As can be seen from the above description, an embodiment of the present invention provides a method for pre-detecting listening test content, including:

[0091] First, pick up the preamble part, the first demarcation label in the listening entity part and / or the postscript part of the listening test content; wherein, the preamble part is the characterization information of the listening test content, the postscript part is the end information of the listening test content, and the first demarcation label in the listening entity part is used to divide the listening entity part into at least two parts, and does not involve the test questions in the listening entity part; perform speech recognition on the preamble part, the first demarcation label and / or the postscript part to generate a speech recognition result; then, detect whether the listening test content is correct according to the speech recognition result, the pre-acquired text content of the preamble part and / or the text content of the postscript part; when the listening test content is correct, detect whether there is any lag in the listening test content according to the speech recognition result.

[0092] In summary, the method for pre-detecting listening test content provided by the present invention, without touching the listening entity part (the actual test questions) of the listening test content, and detecting the listening test content before the test, to eliminate the possibility of test accidents caused by incorrect listening content or problems with the listening test audio file in the prior art, so as to ensure the normal progress of the listening test.

[0093] Regarding step 100, it can be understood that the preamble part is only the characterization information of the listening test content, and does not involve the specific questions in the listening test content. For example, "This is the college entrance examination in 2024", "The listening test officially starts", etc. The postscript part is the end information of the listening test content, such as "XXX listening test ends here", "Please prepare for the written test part of the candidates", etc.

[0094] The first demarcation label is located in the listening entity part, that is, between the preamble part and the postscript part, but does not involve the actual test questions. For example, "Please see the first section of the listening part", "The first section ends here", "The second section ends here", and "The listening part ends here". In this case, the listening test content is not only divided into the preamble part and the postscript part by the second demarcation label (which will be explained later), but also the listening entity part is divided into multiple parts (at least two) by the first demarcation label, and each part involves specific actual test questions. Therefore, the first demarcation label cannot involve the test questions in the listening entity part.

[0095] In a preferred embodiment, a method for pre-detecting the content of a listening exam provided by the present invention is integrated in an electronic device. The electronic device plays and picks up all the content of the listening exam internally, generating only electrical signals without making sounds and without leakage during the playing process. Among them, only the preamble part, the first demarcation label and / or the postscript part are subjected to speech recognition inside the electronic device, and only the preamble part, the first demarcation label and / or the postscript part have the recognition ability.

[0096] Before step 200 is executed, at least one first demarcation label and at least one second demarcation label need to be set in the content of the listening exam. After the speech recognition process recognizes the first demarcation label or the second demarcation label, the speech recognition will no longer be performed (until the next first demarcation label or second demarcation label is recognized). More safely, encrypt the content of the listening exam after (or before, corresponding to the postscript part at this time) the first demarcation label or the second demarcation label so that the speech recognition of the actual content of the listening (exam questions) cannot be performed to prevent the leakage of the listening exam questions.

[0097] The execution process of step 300 is as follows. First, obtain the text content of the preamble part, the first demarcation label in the listening entity part and / or the postscript part (equivalent to the standard content of the preamble part, the first demarcation label in the listening entity part and / or the postscript part) from the exam organizer, and then compare the speech recognition result with this text content. When the two are consistent, it can be considered that the content of the listening exam is correct. Otherwise, it is considered that there is a problem with the storage of the listening exam content, such as an incorrect CD distribution of the listening exam.

[0098] Regarding step 400, it can be understood that even when the content of the listening exam is correct, there may be cases where the listening exam CD has defects resulting in playback jams or complete inability to play. Therefore, it is necessary to continue to detect it. Specifically, it is possible to detect whether there is a jam in the content of the listening exam according to the first time interval between two adjacent sentences in the speech recognition result and a preset first threshold; it is also possible to first determine at least one keyword in the speech recognition result; then, determine the second time interval between the keyword and the initial time of the preamble part and a preset second threshold to detect whether there is a jam in the content of the listening exam. It is also possible to determine at least two keywords in the speech recognition result; and determine the third time interval between the two keywords and a preset third threshold to detect whether there is a jam in the content of the listening exam.

[0099] Based on the idea of step 400, step 200, that is, detecting whether the content of the listening exam is correct, can also be determined by the time interval between sentences, the time interval between the keyword and the initial time of the preamble part, and the time interval between multiple keywords (that is, the above time intervals of each year's exam can be set differently in advance).

[0100] In some embodiments of the present invention, refer to Figure 2 , a pre-detection method for the content of a listening exam further includes:

[0101] Step 500: Set at least one second demarcation label in the content of the listening exam; the second demarcation label is used to divide the preamble part and the listening entity part of the listening exam content, and to divide the listening entity part and the postscript part;

[0102] Specifically, insert a silent segment (zero amplitude) or a specific frequency tone (such as above 20 kHz, inaudible to the human ear) at the position corresponding to the second demarcation label. In a more specific embodiment, the second demarcation label is the spoken word "The listening exam officially starts". When this spoken word is recognized by speech recognition, speech recognition will no longer continue.

[0103] Step 600: Encrypt the listening entity part;

[0104] First, split the audio file of the listening exam content into multiple segments (the part to be encrypted and the part not to be encrypted). Only encrypt the target segment (the listening entity part) (select a symmetric encryption algorithm (such as AES) to encrypt the segmented audio data, and the encryption key and the initial vector need to be saved for decryption. The encryption key and the initial vector are fixedly stored in the playback device designated by the exam organizer). Merge the encrypted segment with the unencrypted part into a complete file. Record the position of the encrypted part through metadata or specific markers for subsequent decryption.

[0105] In some embodiments of the present invention, refer to Figure 3 , step 100 of picking up the preamble part and / or the postscript part of the listening exam content includes:

[0106] Step 101: Determine the position of the second demarcation label;

[0107] When the second demarcation label is a silent segment (zero amplitude) or a specific frequency tone, monitor the position with the following characteristics, which is the position of the second demarcation label:

[0108] Time domain feature: The waveform amplitude continuously remains below a threshold value (such as -50 dB).

[0109] Frequency domain feature: The energy of all frequencies approaches zero.

[0110] Step 102: Pick up the preamble part and / or the postscript part according to the position of the second demarcation label.

[0111] In some embodiments of the present invention, refer to Figure 4 , step 200 includes:

[0112] Step 201: Compare the speech recognition result with the text content;

[0113] It should be noted that the speech recognition result needs to be compared with the text content character by character (comparing each character).

[0114] Step 202: If the speech recognition result is consistent with the text content, determine that the content of the listening exam is correct;

[0115] Step 203: Otherwise, determine that the content of the listening exam is incorrect.

[0116] In some embodiments of the present invention, the text content includes: the second demarcation label (Now is the listening test tone time, the tone test ends here, the listening exam officially starts, and the listening part ends here), the year of the listening exam (This is 2024), the month, and the name (National Unified Examination).

[0117] It should be noted that some of the content in the front or back parts can be replaced. For example, in the 2024 college entrance examination content: the first preposition changes every year. Other front or back parts (prepositions or postpositions) can be reported by voice by calculating the detected time difference, which is used to determine the content correctness and whether there is a lag.

[0118] In some embodiments of the present invention, step 300 includes:

[0119] Detect whether there is a lag in the content of the listening exam according to the first time interval between two adjacent sentences in the speech recognition result and a preset first threshold.

[0120] The method of determining a sentence can be achieved by determining the beginning and end of the sentence. Further, it can be executed through context part-of-speech analysis and grammatical structure integrity in the speech recognition result:

[0121] End-of-sentence feature: The end of a sentence often follows a terminating part of speech (such as the past tense of a verb, sentence-final particles). For example, words like "le", "ma", "ba" in Chinese often appear at the end of a sentence. In this application, the end-of-sentence feature includes common words in the listening exam scenario.

[0122] Beginning-of-sentence feature: The beginning of a sentence may be the subject or a conjunction (such as "however", "therefore").

[0123] Subject-predicate-object integrity: Determine whether the sentence contains a complete subject-predicate-object structure.

[0124] Dependency syntactic tree: Determine whether the sentence components are complete through syntactic analysis.

[0125] In some embodiments of the present invention, see Figure 5 , step 300 includes:

[0126] Step 301: Determine at least one keyword in the speech recognition result;

[0127] Set some non-confidential Chinese content as keywords, and after speech recognition, announce the detection result.

[0128] Step 302: Determine the second time interval between the keyword and the initial time of the preamble part and a preset second threshold to detect whether there is a lag in the listening exam content.

[0129] By detecting and announcing the time difference between the detected keywords and comparing it with the specified time difference, it is possible to know whether there is a lag caused by carrier damage during the playback. It can be understood that it is also possible to detect whether there is a lag in the listening exam content by determining the time interval between the keyword and the end time of the postscript part and a preset threshold.

[0130] Examples of the detection result announcement speech:

[0131] 1. Detected the audio of the XXXX year XXXXXXXX enrollment examination;

[0132] 2. Detected keyword 2, and the time interval from the start detection word is XX milliseconds;

[0133] 3. Detected keyword 3, and the time interval from the start detection word is XX milliseconds;

[0134] …

[0135] n. Detected the end keyword 3, and the time interval from the start detection word is XX milliseconds.

[0136] In some embodiments of the present invention, refer to Figure 6 , Step 300 further includes:

[0137] Step 303: Determine at least two keywords in the speech recognition result;

[0138] Step 304: Determine the third time interval between the two keywords and a preset third threshold to detect whether there is a lag in the listening exam content.

[0139] It can be understood that the preamble part of the listening exam content is mostly in a fixed pattern, that is, the time interval between two keywords does not change much. Therefore, it is possible to determine whether there is a lag in the listening exam content by the time interval between two keywords and the preset range of this time interval.

[0140] In some embodiments of the present invention, step 200 can be implemented in the following manner: First, audio preprocessing (since the audio file instructions for the listening test content generally have a high average price, this step can generally be omitted), feature extraction, model inference, and post-processing are performed:

[0141] Audio preprocessing: Audio files are represented by sound signals (waveforms) and need to be preprocessed to facilitate subsequent processing and analysis.

[0142] Noise reduction: Audio signals usually contain background noise, which can affect the accuracy of recognition. By using noise reduction techniques (such as filters) to reduce background noise, the quality of the speech signal can be improved.

[0143] Framing: Audio signals are continuous, while speech recognition usually analyzes audio within a short time window (e.g., 20 - 40 milliseconds). Therefore, the continuous audio signal needs to be divided into small segments (frames), and these small segments (frames) are called audio frames.

[0144] Window function: After the audio signal is framed, in order to reduce edge effects, each frame is usually windowed (e.g., using a Hamming window or a Hanning window) to make the signal transition smoothly at the boundaries of each frame.

[0145] Feature extraction: Extracting features from audio signals is a key step in speech recognition because the original audio waveform data is too complex to be directly used for recognition. Here, Mel Frequency Cepstral Coefficients (MFCC) are selected as the features for the preface part of the listening test content.

[0146] Mel Frequency Cepstral Coefficients (MFCC): A filter bank based on the Mel frequency scale is used to simulate the auditory characteristics of the human ear. MFCC obtains spectral information by performing a Fourier transform on the audio frames, then performs a Mel transform and a cepstral transform on the spectrum to obtain the coefficients representing the sound features.

[0147] Acoustic model: Used to convert the features extracted from audio signals into probability distributions associated with language units (such as phonemes, syllables, or words). The core goal of the acoustic model is to predict the possible speech units in the audio signal based on the features.

[0148] Language model: Used to calculate the posterior probabilities of the word sequences generated during the speech recognition process to ensure that the recognition results conform to the actual usage rules of the language.

[0149] n-gram Model: Traditional language models, such as unigram (single word), bigram (two-word phrase), and trigram (three-word phrase) models, help determine which phrases are more reasonable by calculating the probability of words appearing in context. The n-gram model calculates the likelihood of a word or phrase by statistically analyzing the joint probability of words in a large corpus.

[0150] Neural Network Language Model (RNN / LSTM): In recent years, language models based on recurrent neural networks or long short-term memory networks have become mainstream. These models can better capture semantic information based on context information and improve the accuracy of speech recognition.

[0151] Decoding and Post-Processing: The decoding process combines the acoustic model and the language model to output the final recognition result. The goal of decoding is to find the most likely sequence of words, that is, to maximize the conditional probability given the features. The decoding process involves the following aspects:

[0152] Viterbi Algorithm: The Viterbi algorithm is the most commonly used dynamic programming algorithm in the decoding process, used for joint decoding between the acoustic model and the language model to calculate the optimal path.

[0153] Beam Search: The beam search algorithm is a heuristic search algorithm. During the decoding process, by retaining several candidate paths with the highest probabilities at each step, it avoids exhaustive search of all possible paths and reduces the computational amount.

[0154] Post-Processing: The decoded result may require some corrections or formatting, such as spelling correction, removal of duplicate words, or repair of grammar errors. In addition, some speech recognition systems may also apply speech context knowledge for further optimization.

[0155] Finally, the decoded output text will provide a written representation of the content spoken in the audio. This is the final result of the speech recognition system, which converts the speech information in the audio into a readable text form.

[0156] In summary, the basic process of speech recognition for the audio file of a listening exam includes:

[0157] 1. Audio Preprocessing: Includes operations such as denoising, framing, and window functions.

[0158] 2. Feature Extraction: By extracting features such as MFCC or spectrogram, the audio signal is converted into numerical features that can represent speech information.

[0159] 3. Acoustic Model: Use HMM or deep learning models (such as DNN, CNN, LSTM, etc.) to predict phonemes, syllables, or words based on features.

[0160] 4. Language model: Use n-gram or neural network language model to constrain the results of speech recognition to conform to the natural rules of language.

[0161] 5. Decoding and post-processing: Optimize the recognition results through Viterbi algorithm, beam search and post-processing and output the final text.

[0162] To further illustrate the solution, the present invention also provides a specific implementation manner of a pre-detection system for the content of a listening exam, which specifically includes the following:

[0163] See Figure 7 , a pre-detection system for the content of a listening exam includes: a sealed and confidential CD signal shielding transmission box (inside which there is a CD player (or USB player and other playback tools), a speech recognition and speech generation processor, a speaker, a text comparison module, and an open cover detection and time recording module.

[0164] The content included in the audio file of the listening exam broadcast CD is divided into a confidential part and a non-confidential part. The Chinese content is the description of the exam and is basically fixed in content and does not require confidentiality, such as "Here in 2024 is the college entrance examination for ordinary higher education", "The listening exam officially starts", etc. The above CD player is internally fixed with the key and initial vector corresponding to the encryption of the audio file of the listening exam broadcast CD.

[0165] The playback tool plays the audio signal in the carrier, and the generated audio electrical signal is sent to the speech recognition module. Some non-confidential Chinese content is set as the detection keyword. After being recognized by the speech recognition module, the detection result is announced through the speaker. By detecting and comparing the time difference between the detected keywords with the specified time difference, it is possible to know whether there is a freeze caused by carrier damage during the playback process.

[0166] In addition, the above pre-detection system for the content of a listening exam has a physical and electromagnetic leakage protection device. The corresponding physical device is used for protection, and the corresponding electromagnetic leakage protection device has the following functions:

[0167] Reading device protection: Use an anti-electromagnetic leakage optical drive, or install a metal shielding cover (such as copper mesh + conductive gasket) for an ordinary optical drive. It is prohibited to use wireless transmission functions such as Wi-Fi and Bluetooth in the reading environment to prevent data from leaking through electromagnetic wave bypass.

[0168] Signal interference:

[0169] Active interference: Deploy a broadband noise jammer in the CD reading area to cover the working frequency band of the optical drive (such as 1 - 5 GHz) to interfere with potential eavesdropping devices.

[0170] Power filtering: Install power line filters for the optical drive and the host to suppress conductive electromagnetic leakage.

[0171] Anti-laser detection:

[0172] Anti-peeping film layer: Coat a light-scattering coating (such as a titanium dioxide nanoparticle film) on the surface of the optical disc, so that the laser probe cannot directly read the pit / land structure.

[0173] Dynamic encryption: Use an optical disc that needs to cooperate with a dedicated hardware key, and real-time decryption is required during reading to avoid the direct exposure of plaintext data by electromagnetic signals.

[0174] As can be seen from the above description, the specific implementation of the present invention provides a method for pre-detecting the content of a listening test, including: First, pick up the preface part of the listening test content; where the preface part is the characterization information of the listening test content; and perform speech recognition on the preface part to generate a speech recognition result; Then, detect whether the listening test content is correct according to the speech recognition result and the pre-acquired text content of the preface part; Finally, if the listening test content is correct, detect whether there is any lag in the listening test content according to the speech recognition result.

[0175] In summary, the method for pre-detecting the content of a listening test provided by the present invention, without contacting the listening entity part (the real test questions) of the listening test content, and detecting the listening test content before the test, so as to eliminate the possibility of test accidents caused by errors in the listening content or problems with the listening test audio file in the prior art, and ensure the normal progress of the listening test.

[0176] Based on the same inventive concept, the embodiments of the present application also provide a device for pre-detecting the content of a listening test, which can be used to implement the method described in the above embodiments, as in the following embodiments. Since the principle of solving problems by the device for pre-detecting the content of a listening test is similar to that of the method for pre-detecting the content of a listening test, the implementation of the device for pre-detecting the content of a listening test can refer to the implementation of the method for pre-detecting the content of a listening test, and the repeated parts will not be elaborated. Hereinafter, the term "unit" or "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0177] The embodiments of the present invention provide a specific implementation of a device for pre-detecting the content of a listening test that can implement the method for pre-detecting the content of a listening test. See Figure 8 , a device for pre-detecting the content of a listening test specifically includes the following:

[0178] The pickup module 10 is used to pick up the preamble part, the first demarcation label in the listening entity part and / or the postscript part of the listening test content; wherein, the preamble part is the characterization information of the listening test content, the postscript part is the end information of the listening test content, and the first demarcation label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the test questions in the listening entity part;

[0179] The speech recognition result generation module 20 is used to perform speech recognition on the preamble part, the first demarcation label and / or the postscript part to generate a speech recognition result;

[0180] The listening content correctness detection module 30 is used to detect whether the listening test content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first demarcation label and / or the text content of the postscript part;

[0181] The listening content jitter detection module 40 is used to detect whether there is jitter in the listening test content according to the speech recognition result when the listening test content is correct.

[0182] In some embodiments of the present invention, refer to Figure 9 , a pre-detection device for listening test content, further comprising:

[0183] The demarcation label setting module 50 is used to set at least one second demarcation label in the listening test content; the second demarcation label is used to divide the preamble part from the listening entity part of the listening test content, and to divide the listening entity part from the postscript part;

[0184] The listening entity part encryption module 60 is used to encrypt the listening entity part;

[0185] Refer to Figure 10 , the preamble part pickup module 10 includes:

[0186] The demarcation label position determination unit 10a is used to determine the position of the second demarcation label;

[0187] The preamble part pickup unit 10b is used to pick up the preamble part and / or the postscript part according to the position of the second demarcation label.

[0188] In some embodiments of the present invention, refer to Figure 11 , the listening content correctness detection module 30 includes:

[0189] The comparison unit 30a is used to compare the speech recognition result with the text content;

[0190] The examination content correct judgment unit 30b is used to judge that the listening examination content is correct when the speech recognition result is consistent with the text content;

[0191] The examination content wrong judgment unit 30c is used to otherwise judge that the listening examination content is wrong.

[0192] In some embodiments of the present invention, the text content includes: the first demarcation label, the second demarcation label, the year, month and name of the listening examination.

[0193] In some embodiments of the present invention, the listening content stutter detection module includes:

[0194] The first unit for detecting stutter in the listening content is used to detect whether there is a stutter in the listening examination content according to the first time interval between two adjacent sentences in the speech recognition result and a preset first threshold.

[0195] In some embodiments of the present invention, see Figure 12 , the listening content stutter detection module 40 includes:

[0196] The first unit 40a for determining keywords is used to determine at least one keyword in the speech recognition result;

[0197] The second unit 40b for detecting stutter in the listening content is used to detect whether there is a stutter in the listening examination content according to the second time interval between the keyword and the initial time of the preface part and a preset second threshold.

[0198] In some embodiments of the present invention, see Figure 13 , the listening content stutter detection module 40 further includes:

[0199] The second unit 40c for determining keywords is used to determine at least two keywords in the speech recognition result;

[0200] The third unit 40d for detecting stutter in the listening content is used to detect whether there is a stutter in the listening examination content according to the third time interval between the two keywords and a preset third threshold.

[0201] As can be seen from the above description, an embodiment of the present invention provides a pre-detection device for the content of a listening exam, including: a preamble picking module for picking up the preamble of the listening exam content; wherein the preamble is the characterization information of the listening exam content; a speech recognition result generation module for performing speech recognition on the preamble to generate a speech recognition result; a listening content correctness detection module for detecting whether the listening exam content is correct according to the speech recognition result and the pre-acquired text content of the preamble; a listening content stutter detection module for detecting whether there is a stutter in the listening exam content according to the speech recognition result if the listening exam content is correct.

[0202] In summary, without touching the physical part of the listening exam content (the actual exam questions), the present invention detects the listening exam content before the exam to eliminate the possibility of exam accidents caused by incorrect listening content or problems with the listening exam audio file in the prior art, and ensure the normal progress of the listening exam.

[0203] An embodiment of the present application also provides a specific implementation manner of an electronic device that can implement all steps in the pre-detection method for the content of a listening exam in the above embodiment. Refer to Figure 14 , and the electronic device specifically includes the following contents:

[0204] A processor (processor) 1201, a memory (memory) 1202, a communication interface (Communications Interface) 1203, and a bus 1204;

[0205] Among them, the processor 1201, the memory 1202, and the communication interface 1203 complete mutual communication through the bus 1204; the communication interface 1203 is used to implement information transmission between related devices such as the server-side device and the user-side device;

[0206] The processor 1201 is used to call the computer program in the memory 1202. When the processor executes the computer program, all steps in the pre-detection method for the content of a listening exam in the above embodiment are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0207] Step 100: Pick up the preamble of the listening exam content, the first delimiter label and / or the postscript in the listening entity part; wherein the preamble is the characterization information of the listening exam content, the postscript is the end information of the listening exam content, and the first delimiter label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the exam questions in the listening entity part;

[0208] Step 200: Perform speech recognition on the preamble part, the first delimiter label, and / or the postscript part to generate a speech recognition result;

[0209] Step 300: Detect whether the listening test content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first delimiter label, and / or the text content of the postscript part;

[0210] Step 400: If the listening test content is correct, detect whether there is any lag in the listening test content according to the speech recognition result.

[0211] An embodiment of the present application further provides a computer-readable storage medium capable of implementing all steps in the above-mentioned pre-detection method for listening test content. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, all steps in the above-mentioned pre-detection method for listening test content are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0212] Step 100: Pick up the preamble part of the listening test content, the first delimiter label in the listening entity part, and / or the postscript part; wherein, the preamble part is the characterization information of the listening test content, the postscript part is the end information of the listening test content, and the first delimiter label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the question content in the listening entity part;

[0213] Step 200: Perform speech recognition on the preamble part, the first delimiter label, and / or the postscript part to generate a speech recognition result;

[0214] Step 300: Detect whether the listening test content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first delimiter label, and / or the text content of the postscript part;

[0215] Step 400: If the listening test content is correct, detect whether there is any lag in the listening test content according to the speech recognition result.

[0216] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the hardware + program type embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0217] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0218] Although this application provides method operation steps such as in the embodiments or flowcharts, based on routine or non-creative labor, there may be more or fewer operation steps. The order of steps listed in the embodiments is only one way among many possible execution orders and does not represent the only execution order. When the actual device or client product is executed, it can be executed in the order shown in the embodiments or the figures or in parallel (e.g., in an environment with parallel processors or multi-threaded processing).

[0219] For convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing the embodiments of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0220] Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and the structures within the hardware component.

[0221] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0222] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM). The memory is an example of computer-readable media.

[0223] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment. In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of this specification. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0224] The above are only the embodiments of the embodiments of this specification and are not used to limit the embodiments of this specification. For those skilled in the art, various changes and modifications can be made to the embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of this specification shall be included within the scope of the claims of the embodiments of this specification.

Claims

1. A pre-detection method for the content of a listening exam, characterized in that, Including: Picking up the preamble part of the listening test content, the first demarcation label in the listening entity part and / or the postscript part; wherein, the preamble part is the characterization information of the listening test content, the postscript part is the ending information of the listening test content, and the first demarcation label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the examination questions in the listening entity part; Performing speech recognition on the preamble part, the first demarcation label and / or the postscript part to generate a speech recognition result; Detecting whether the listening test content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first demarcation label and / or the text content of the postscript part; If the listening test content is correct, detecting whether there is any lag in the listening test content according to the speech recognition result.

2. The pre-detection method according to claim 1, characterized in that Also including: Setting at least one second demarcation label in the listening test content; The second demarcation label is used to divide the preamble part from the listening entity part of the listening test content and to divide the listening entity part from the postscript part; Encrypting the listening entity part; The picking up the preamble part and / or the postscript part of the listening test content includes: Determining the position of the second demarcation label; Picking up the preamble part and / or the postscript part according to the position of the second demarcation label.

3. The pre-detection method according to claim 1, wherein The detecting whether the listening test content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first demarcation label and / or the text content of the postscript part includes: Comparing the speech recognition result with the text content; If the speech recognition result is consistent with the text content, determining that the listening test content is correct; Otherwise, determining that the listening test content is incorrect.

4. The pre-detection method according to claim 2, characterized in that, The text content includes: the first demarcation label, the second demarcation label, the year, month and name of the listening test.

5. The pre-detection method according to claim 1, wherein The detecting whether there is any lag in the listening test content according to the speech recognition result includes: Detecting whether there is any lag in the listening test content according to the first time interval between two adjacent sentences in the speech recognition result and a preset first threshold.

6. The pre-detection method according to claim 1, characterized in that The detecting whether there is any lag in the listening test content according to the speech recognition result includes: Determining at least one keyword in the speech recognition result; Determining whether there is any lag in the listening test content according to the second time interval between the keyword and the initial time of the preamble part and a preset second threshold.

7. The pre-detection method according to claim 6, characterized in that, The detecting whether there is any lag in the listening test content according to the speech recognition result further includes: Determining at least two keywords in the speech recognition result; Determining whether there is any lag in the listening test content according to the third time interval between the two keywords and a preset third threshold.

8. A pre-detection device for the content of a listening test, characterized in that, Including: A picking module, configured to pick up the preamble part, the first demarcation label in the listening entity part and / or the postscript part of the listening exam content; wherein, the preamble part is the characterization information of the listening exam content, the postscript part is the ending information of the listening exam content, and the first demarcation label in the listening entity part is used to divide the listening entity part into at least two parts and does not involve the exam questions in the listening entity part; A speech recognition result generation module, configured to perform speech recognition on the preamble part, the first demarcation label and / or the postscript part to generate a speech recognition result; A listening content correctness detection module, configured to detect whether the listening exam content is correct according to the speech recognition result, the pre-acquired text content of the preamble part, the text content of the first demarcation label and / or the text content of the postscript part; A listening content stutter detection module, configured to detect whether there is a stutter in the listening exam content according to the speech recognition result if the listening exam content is correct.

9. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the pre-detection method for the listening exam content according to any one of claims 1 to 7 are implemented.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, the steps of the pre-detection method for the listening exam content according to any one of claims 1 to 7 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the pre-detection method for the listening exam content according to any one of claims 1 to 7 are implemented.