Broadcast test method and device, wearable device, storage medium
By building a broadcast test model, using a multi-head attention mechanism and voice output module, combined with regular expressions and semantic feature library, the problem of insufficient broadcast test in the existing technology is solved, and more accurate and reliable broadcast test is achieved.
Patent Information
- Application Number
- CN202510322862.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing broadcast testing methods mainly focus on functional testing, and it is difficult to effectively evaluate the quality and reliability of voice broadcasts, resulting in insufficient broadcast testing.
Build a broadcast test model, and use a multi-head attention mechanism and speech output module, combined with regular expressions and semantic feature library to identify and adjust form and semantic errors to realize data-driven testing.
Improve the accuracy and reliability of broadcast testing, ensure the stability and consistency of the model, and timely discover and correct errors, and improve broadcast quality.
Smart Images

Figure CN119851696B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of data testing, and more specifically, relates to a method and apparatus for broadcast testing, a wearable device, and a storage medium. Background Art
[0002] With the rapid development of information technology, voice broadcast systems have been widely used in many fields. In a smart home environment, devices such as smart speakers can broadcast news, weather conditions, reminder items, etc., providing users with an intelligent home interaction experience; in mobile Internet applications, various reading software, learning software, etc. have also introduced voice broadcast functions, enabling users to obtain information when they are unable to read the screen content, such as blind assistive reading applications.
[0003] To ensure the quality and reliability of voice broadcasts, testing is crucial. Traditional testing methods mainly focus on functional testing, that is, checking whether the system can normally start voice broadcasts, whether it can accurately identify input text information, and other basic functions. This testing method can, to a certain extent, discover some obvious defects, but there are deficiencies in the testing of broadcast quality.
[0004] Therefore, there is an urgent need for an accurate and reliable broadcast testing method. Summary of the Invention
[0005] The purpose of the present disclosure is to provide a method and apparatus for broadcast testing, a wearable device, and a storage medium to improve the accuracy and reliability of broadcast testing.
[0006] In the first aspect of the embodiments of the present disclosure, a method for broadcast testing is provided, including:
[0007] Constructing a broadcast testing model;
[0008] Inputting a plurality of test information into the broadcast testing model to determine a plurality of target broadcast information, where the plurality of test information is information in a test set, and the test set includes a plurality of test information and the corresponding broadcast information for each test information;
[0009] In response to the number of error information in the plurality of target broadcast information being greater than a first preset value and the type of the error information being a formal error, adjusting the broadcast testing model based on a first method; and performing the step of inputting the plurality of test information into the broadcast testing model to determine the plurality of target broadcast information until the number of error information in the plurality of target broadcast information is less than or equal to the first preset value;
[0010] In response to the number of error messages in multiple target broadcast messages being greater than a first preset value and the type of the error messages being semantic errors, adjust the broadcast test model based on a second method; and perform the step of inputting multiple test messages into the broadcast test model to determine multiple target broadcast messages until the number of error messages in the multiple target broadcast messages is less than or equal to the first preset value;
[0011] Wherein, the processing methods of the first method and the second method are different.
[0012] In a second aspect of the embodiments of the present disclosure, there is provided a broadcast test device, including:
[0013] A model construction module, configured to construct a broadcast test model;
[0014] A model test module, configured to input multiple test messages into the broadcast test model to determine multiple target broadcast messages, where the multiple test messages are messages in a test set, and the test set includes multiple test messages and the corresponding broadcast messages for each test message;
[0015] A model adjustment module, configured to, in response to the number of error messages in multiple target broadcast messages being greater than a first preset value and the type of the error messages being formal errors, adjust the broadcast test model based on a first method; and perform the step of inputting multiple test messages into the broadcast test model to determine multiple target broadcast messages until the number of error messages in the multiple target broadcast messages is less than or equal to the first preset value;
[0016] In response to the number of error messages in multiple target broadcast messages being greater than a first preset value and the type of the error messages being semantic errors, adjust the broadcast test model based on a second method; and perform the step of inputting multiple test messages into the broadcast test model to determine multiple target broadcast messages until the number of error messages in the multiple target broadcast messages is less than or equal to the first preset value;
[0017] Wherein, the processing methods of the first method and the second method are different.
[0018] In a third aspect of the embodiments of the present disclosure, there is provided a wearable device, including a memory, a processor, and a computer program stored in the memory and running on the processor, and when the processor executes the computer program, the steps of the above-mentioned broadcast test method are implemented.
[0019] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned broadcast test method are implemented.
[0020] The beneficial effects of the broadcast test method, device, wearable device, and storage medium provided by the embodiments of the present disclosure are as follows:
[0021] The present disclosure realizes data-driven testing by inputting multiple test messages into a model and determining corresponding target broadcast messages, making the testing more comprehensive and systematic. The test set includes multiple test messages and corresponding broadcast messages, which can cover more test scenarios. In the present disclosure, when there are errors in the broadcast messages generated by the model, they can be discovered and corrected in a timely manner, which helps to continuously optimize and improve the model, and improve the accuracy and reliability of the model. The present disclosure ensures the stability and consistency of the model by repeatedly executing the steps of inputting test messages into the model and checking the target broadcast messages until there are no error messages. The process of iterative testing helps to gradually eliminate potential problems in the model, improve the broadcast quality of the model, and enhance the accuracy and reliability of broadcast testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of a broadcast testing method provided by an embodiment of the present disclosure;
[0024] Figure 2 It is a structural block diagram of a broadcast testing device provided by an embodiment of the present disclosure;
[0025] Figure 3 It is a schematic block diagram of a wearable device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0027] To make the purpose, technical solutions, and advantages of the present disclosure clearer, the following will be illustrated through specific embodiments in conjunction with the drawings.
[0028] Please refer to Figure 1 , Figure 1 which is a flowchart of a broadcast testing method provided by an embodiment of the present disclosure. The method includes:
[0029] S101: Construct a broadcast testing model.
[0030] In this embodiment, the process of constructing the broadcast test model may include:
[0031] Determine the basic architecture of the model;
[0032] Construct a language processing module;
[0033] Construct a voice output module;
[0034] Determine the broadcast test model based on the basic architecture, the language processing module, and the voice output module.
[0035] In this embodiment, the goal of the broadcast test is to test and adjust a model that broadcasts the received content in an actual tone. The basic architecture of the broadcast test model can be based on the Transformer architecture, or a recurrent neural network and its variants, etc. The language processing module can be based on the combination of the multi-head attention mechanism and the feed-forward neural network. The feed-forward neural network performs further non-linear transformation on the features processed by the attention mechanism to enhance the model's ability to express complex semantics, enabling it to learn deeper language structures and semantic relationships, such as understanding language features such as modification relationships, subject-predicate-object structures in sentences that affect the broadcast content.
[0036] The multi-head attention mechanism can parallelly capture the relationships between words in the text from multiple perspectives and extract semantic features. For example, when processing the sentence "Today's weather is suitable for going out to play", different heads may respectively focus on different aspects of semantic information such as the relevance between "weather" and "play", and the overall sentiment tendency of the sentence.
[0037] The multi-head attention mechanism can also parallelly process the influence of multiple influencing factors on the broadcast information. Each attention mechanism in the multi-head attention mechanism respectively focuses on the correlation degree of different influencing factors on the broadcast information and finally outputs a broadcast information.
[0038] The voice output module can convert the received text information or the processed voice-related data into sound signals and play them through the corresponding sound-producing device, thereby achieving the effect of voice output and completing the "speaking" link in the entire voice interaction process. The voice output module can broadcast according to the broadcast information output by the voice processing module. The voice output module can convert the text into corresponding voice waveform data through pre-set voice synthesis rules or based on a trained voice synthesis model.
[0039] S102: Input multiple test messages into the broadcast test model to determine multiple target broadcast messages. The multiple test messages are the messages in the test set. The test set includes multiple test messages and the corresponding broadcast messages for each test message.
[0040] In this embodiment, the test information is a set of input text contents specially prepared for testing the voice broadcast test model. These contents are carefully selected and organized to cover various possible text situations, language expression forms, application scenarios, etc., so as to comprehensively examine the voice broadcast ability of the model. The test information can include various types of information such as text information, picture information, website information, compressed package information, etc.
[0041] The test set is a set composed of a large amount of test information and the corresponding correct voice broadcast information. It is the basis for the entire test process. By comparing the target voice broadcast information output by the model with the correct voice broadcast information preset in the test set, it can be determined whether the model makes mistakes and what aspects need to be adjusted.
[0042] For example, the test set can be presented in the form of a table, as shown in Table 1:
[0043] Table 1 Test Set
[0044]
[0045] It should be noted that the content in parentheses in the table does not belong to the voice broadcast content, but is a limitation on the tone of the voice broadcast. That is, the voice broadcast information in the test set not only includes the text of the voice broadcast, but also includes the corresponding tone. The form in which the voice broadcast information in the test set exists can be a piece of voice. For better display and understanding, it is presented in the form of text in Table 1.
[0046] Similarly, the "picture type" in the test information is not the text "picture type", but the test information is a picture. At this time, the corresponding voice broadcast information output is "The other party sent a picture". That is, the text in double quotes in Table 1 is the actual input text and the actual output text.
[0047] The target voice broadcast information refers to the voice broadcast information output after inputting the test information into the voice broadcast test model.
[0048] S103: In response to the number of error messages in the multiple target voice broadcast information being greater than the first preset value, and the type of the error message being a form error, adjust the voice broadcast test model based on the first method; and, execute the step of inputting the multiple test information into the voice broadcast test model to determine the multiple target voice broadcast information until the number of error messages in the multiple target voice broadcast information is less than or equal to the first preset value;
[0049] In response to the number of error messages in multiple target broadcast messages being greater than a first preset value and the type of the error message being a semantic error, adjust the broadcast test model based on a second method; and perform the step of inputting multiple test messages into the broadcast test model to determine multiple target broadcast messages until the number of error messages in the multiple target broadcast messages is less than or equal to the first preset value;
[0050] Among them, the processing methods of the first method and the second method are different.
[0051] In this embodiment, an error message refers to the content corresponding to the places where the target broadcast message does not meet the requirements or is inconsistent compared with the corresponding correct broadcast message in the test set. The types of error messages may be diverse, such as content deviation caused by semantic understanding errors, the tone of the broadcast not meeting the expectations, quality problems such as noise in speech synthesis, etc. The first preset value can be a value related to the proportion of the number of target broadcast messages. For example, the first preset value can be a preset proportion of the number of target broadcast messages. Specifically, the preset proportion can be 0.5%, and the preset proportion can be set according to the usage scenario or the preference of the tester.
[0052] Whether there are error messages in multiple target broadcast messages can be calculated by calculating the matching degree between the target broadcast message and the corresponding correct broadcast message in the test set.
[0053] In this embodiment, in response to the similarity between the target broadcast message and the broadcast message in the test set being less than or equal to a first similarity threshold, define the target broadcast message as an error message.
[0054] The similarity can be calculated by respectively extracting the Mel-frequency cepstral coefficients of the target broadcast message and the broadcast message and calculating the cosine similarity. The first similarity threshold can be determined through experiments.
[0055] For example, the correct broadcast message corresponding to "Tomorrow is a holiday" in the test set is broadcast in a cheerful tone, while the target broadcast message output by the model is "Tomorrow is a holiday" broadcast in a flat and emotionless tone. Here, the semantic requirement is an error message; or the model outputs "I'm going to the supermarket to buy apples and bananas" as "I'm going to the store to buy apples and", which is an error message of missing words. Or the model outputs "www.xxxxxx.com" as "w, w, w, dot, xxxxxx", that is, the website link is broadcast in order, which belongs to a form error.
[0056] In this embodiment, the model can be adjusted according to different error types, and the broadcast test model can be adjusted by adjusting the broadcast feature library or activation function of the broadcast test model, etc.
[0057] In this embodiment, the first method can be to process and express error information based on regular expressions, and the second method can be to adjust the broadcast feature library according to the semantic features in the error information. This is because considering the essential differences in the types of error information, for example, formal errors mainly involve that the information does not meet the requirements in terms of its manifestation form, such as format errors, character errors, punctuation errors, etc., which have certain regularity and patterns. Regular expressions have powerful pattern matching and text processing capabilities, and can effectively identify and process such errors with obvious pattern features. Therefore, processing error information based on regular expressions can specifically search for, correct, and standardize formal errors, thereby adjusting the broadcast feature library and improving the accuracy of the broadcast test model in terms of form.
[0058] Semantic errors, on the other hand, refer to deviations in the meaning and content understanding of information, and the problem lies in the understanding and expression of the semantics of the information. It does not have obvious external patterns like formal errors, but requires in-depth analysis at the connotation level of the information. Therefore, it is necessary to determine the semantic features in the error information, such as analyzing the semantic relationships between words and the logical structures of sentences in the error information, and adjusting the broadcast feature library from the semantic level to improve the broadcast test model's ability to understand and process semantics.
[0059] As can be seen from the above, the present disclosure realizes data-driven testing by inputting multiple test information into the model and determining the corresponding target broadcast information, making the testing more comprehensive and systematic. The test set includes multiple test information and corresponding broadcast information, which can cover more test scenarios. In the present disclosure, when there are errors in the broadcast information generated by the model, they can be discovered and corrected in a timely manner, which helps to continuously optimize and improve the model, and improve the accuracy and reliability of the model. The present disclosure ensures the stability and consistency of the model by repeatedly executing the steps of inputting test information into the model and checking the target broadcast information until there is no error information. The process of iterative testing helps to gradually eliminate potential problems in the model, improve the broadcast quality of the model, and improve the accuracy and reliability of the broadcast test.
[0060] In an embodiment of the present disclosure, the broadcast test model includes a broadcast feature library;
[0061] Adjusting the broadcast test model based on the first method includes:
[0062] Processing error information based on regular expressions, and adjusting the broadcast feature library based on the processed error information;
[0063] Adjusting the broadcast test model based on the second method includes:
[0064] Determining the semantic features in the error information, and adjusting the broadcast feature library based on the semantic features.
[0065] In this embodiment, the broadcast feature library is an aggregate in the broadcast test model for storing various types of broadcast-related feature information. These feature information can cover multiple aspects, such as the broadcast tone rules corresponding to different language expression forms, the appropriate speech rate and volume settings in different scenarios, the semantic understanding key points associated with various vocabulary or sentence structures, etc. They provide the basis and reference for the model to generate accurate and reasonable broadcast content, and help the model make appropriate broadcast decisions when facing different input information.
[0066] Or, it is a library that stores various types of information and their corresponding broadcast forms.
[0067] In this embodiment, error information can be input into the first neural network model to obtain the error type of the error information. The first neural network model is obtained through a dataset composed of a certain number of error information and their corresponding error types. A certain number means a number sufficient for the first neural network model to be trained and tested, which can be determined based on experience or the number of data in the dataset when solving related problems in this field.
[0068] For example, when the test information is a website address, the corresponding broadcast form should be "The other party sent a website address", or when the test information is a compressed package, the corresponding broadcast form should be "The other party sent a compressed package".
[0069] The error type may be a text-type error or a non-text form error. For example, an incorrect expression of an emoji.
[0070] For example, the meaning of the emoji [smile] is smile, friendly, etc., and the incorrect expression is sad, heartbroken, etc.
[0071] Regular expressions describe the characteristics of text by defining specific pattern rules. Processing error information based on regular expressions is to use a pre-set regular expression pattern that meets the requirements of corresponding error form correction to identify the text part in the error information that conforms to this pattern, and then perform operations such as replacement and correction on it according to the rules to make its form correct and standardized, so that it can be better processed by the model and used to adjust the broadcast feature library later.
[0072] For example, if it is a situation where punctuation marks are used incorrectly, such as multiple consecutive commas in a sentence being used wrongly (e.g., "Today, I, go to the park, play"), the regular expression: / ,{2,} / can be used to match the situation where two or more consecutive commas appear, and then replace them with a single comma to make the sentence punctuation more standardized, becoming "Today, I go to the park to play".
[0073] Or express the website URL and the compressed package through regular expressions. For example, the basic website URL format is protocol: / / domain name / path. Specifically, https: / / www.example.com / path / to / page.html. A simple regular expression to match such a website URL can be: ^(https?: / / )?([a-zA-Z0-9.-]+)(:\d+)?( / [a-zA-Z0-9_. / ?%&=]*)?$.
[0074] ^(https?: / / )? : ^ means to match the start of the string. https? means http or https. : / / is the fixed part after the protocol.? means that the previous https: / / is optional because some website URLs may not write the protocol part, and the browser will default to using http or https.
[0075] ([a-zA-Z0-9.-]+) : This is a capture group used to match the domain name part. It can contain letters, numbers, dots, and hyphens. + means it appears at least once. For example, www.example.com.
[0076] (:\d+)? : This part is optional and is used to match the port number. : is the prefix of the port number. \d+ means to match one or more digits. For example, :8080.
[0077] ( / [a-zA-Z0-9_. / ?%&=]*)?$ : / is the start part of the path. [a-zA-Z0-9_. / ?%&=]* means the path can contain letters, numbers, underscores, dots, slashes, question marks, percent signs, ampersands, and equal signs. * means it can appear zero or more times. $ means to match the end of the string.
[0078] The error-prone expression forms of the model can be expressed through similar regular expressions respectively, and the expressed expressions and the corresponding correct expression forms are stored in the broadcast feature library.
[0079] When the error type is a semantic error, the semantic features in the error message can be extracted and the corresponding correct semantics are stored in the broadcast feature library.
[0080] Semantic features can refer to the keywords representing semantics in the text. For example, for a sentence: "Have you eaten today?" The semantic features in this sentence can be "ma" and "?", that is, the features representing tone and semantics. If the output of the broadcast test model is the very plain "Have you eaten today.", it belongs to a semantic expression error. At this time, "ma", "?", and the expressions representing interrogation and rising tone can be added to the broadcast feature library.
[0081] From the above, it can be concluded that this embodiment can handle different types of broadcast errors in a more targeted manner by distinguishing between formal errors and semantic errors. Formal errors mainly focus on the surface features of the text, such as punctuation, format, etc., while semantic errors go deep into the meaning level of the text, such as tone, contextual understanding, etc., which improves the efficiency and accuracy of the model in handling errors. The present disclosure uses regular expressions to identify and correct error messages, and has a high degree of flexibility, and can customize corresponding regular expression patterns according to different error types. Through adjustments based on error information, the broadcast feature library of the present disclosure can continuously absorb new feature information, including correct broadcast forms, semantic features, etc., thereby continuously improving the broadcast capability and adaptability of the model, and improving the accuracy and reliability of the broadcast test.
[0082] In one embodiment of the present disclosure, determining semantic features in error information and adjusting a report feature library based on the semantic features includes:
[0083] In response to the semantic error being an error in the semantic expression of the text, extracting semantic features of the text, and adding the semantic features to a broadcast feature library;
[0084] In response to the semantic error being a non-text semantic expression error, initial semantics and correct semantics of the non-text are acquired, and the initial semantics and correct semantics are added to the announcement feature library.
[0085] In this embodiment, a text semantic expression error may be a mismatch in tone. For example, in the sentence “I am particularly happy today!”, there are “especially”, “happy” and “!”, which indicate a strong tone and express excitement, while the output of the broadcast test model is a bland or even sad “I am particularly happy today.” In this case, it is determined to be a semantic expression error.
[0086] Errors in text semantic expression refer to problems in the meaning conveyed by the text itself. This is mainly due to factors such as the text's word choice, grammatical structure, and logical relationships, which cause the semantics expressed to not conform to normal language expression rules or to be inconsistent with the intended meaning.
[0087] Semantic features in test information can be extracted based on word vector methods. Semantic features are not limited to text, but can also be vector expressions corresponding to text that represents semantics. The same or similar semantics can be stored as the same vector in the report feature library. When the same or similar semantic features appear again in the test information, the correct semantics in the report feature library can be directly called for reporting.
[0088] When the semantic error is not a non - text semantic expression error, the initial and correct non - text semantics can be extracted. Taking emoticons as an example, the initial semantics of an emoticon refers to the meaning defined by the designer when creating the emoticon. For example, for the emoticon [smile], its initial meaning is [smile]. However, due to the public's understanding and certain characteristics of the emoticon, when the public uses this emoticon, the meaning it conveys may be [I'm convinced], [speechless], which is quite different from the initial meaning. At this time, the correct meaning of the emoticon is different from the initial meaning.
[0089] In this embodiment, it is also possible to: in response to the semantic error being a non - text semantic expression error, obtain the initial and correct non - text semantics, and obtain the semantic features of the text before and after the non - text, and add the initial semantics, correct semantics, and semantic features to the broadcast feature library.
[0090] Whether the correct meaning represented by an emoticon is different from the initial meaning can be judged based on the semantic features of the text before and after the emoticon. For example, in the sentences "Had a great time today, thank you [smile]" and "You were too excessive today [smile]", the meanings represented by the [smile] emoticon in the two messages are "thank you" and "speechless" respectively. The meaning of the emoticon in the first sentence is the same as the initial meaning, while the meaning of the emoticon in the second sentence is different from the initial meaning. The semantic features in the first sentence can be "very", "happy", and "thank you", and the semantic feature in the second sentence can be "excessive". "Excessive" and "speechless" can be stored in the broadcast feature library. It should be noted that "excessive" is the semantic feature in the sentence, and "speechless" is the meaning of the emoticon. The storage form of "excessive" in the broadcast feature library can be a vector or text.
[0091] When the word "excessive" appears again in a sentence, the broadcast test model can directly express the [smile] emoticon as "speechless" according to the information stored in the broadcast feature library. In the process of continuous training and extraction, more semantic features and their corresponding emoticon meanings can be stored in the broadcast feature library.
[0092] It can be concluded from the above that by extracting the semantic features in the text and adding them to the broadcast feature library, the present disclosure can enhance the model's ability to understand text semantics, help the model more accurately capture key information such as emotions and tones in the text, thereby improving the accuracy and naturalness of the broadcast. By obtaining the initial and correct non - text semantics and associating them with the semantic features of the text before and after, the present disclosure improves the model's ability to understand and recognize non - text information, helps the model make more appropriate broadcast decisions in a complex and changing communication environment, and improves the accuracy and reliability of the broadcast test.
[0093] In one embodiment of the present disclosure, a plurality of test information is input into a broadcast test model to determine a plurality of target broadcast information, including:
[0094] Multiple test information are input into the language processing module, and the multiple test information are processed based on the multi-head attention mechanism to obtain multiple broadcast information.
[0095] Based on the multi-head attention mechanism, multiple test information is processed to obtain multiple broadcast information;
[0096] The language processing module is a module in the broadcast test model; the language processing module is built based on the multi-head attention mechanism.
[0097] In this embodiment, the test information is a set of input text contents specially prepared for testing the broadcast test model. They are carefully selected and organized, covering various possible language expressions, different application scenarios and rich semantic situations, etc., with the purpose of comprehensively examining the broadcasting ability of the model.
[0098] The language processing module is responsible for processing the input text information at the language level. Its core construction foundation is the multi-head attention mechanism. Through the powerful feature capture and semantic understanding capabilities of the multi-head attention mechanism, it can deeply analyze the input text and explore the relationship between elements in different positions.
[0099] The multi-head attention mechanism comprehensively considers information at different positions by calculating multiple "heads" in parallel. It can capture important features such as semantic relationships and long-distance dependencies in text sequences from multiple angles at the same time.
[0100] Take the sentence "I like to read on sunny afternoons" as an example. One "head" may focus on the grammatical structure relationship between words in the sentence, judging the subject, predicate, object and other components; another "head" may focus more on analyzing the semantic association between words, such as the connection between "sunny" and "afternoon" and "reading" in a pleasant atmosphere; and another "head" may pay attention to the overall emotional tendency of the text, which is positive and pleasant. By working in parallel with multiple "heads" like this, all aspects of the sentence can be mined to help the language processing module better understand the text.
[0101] As can be seen from the above, by introducing the multi-head attention mechanism, the language processing module can analyze the input text information more deeply. It not only considers the direct connections between various elements in the text, but also captures complex semantic relationships and long-distance dependencies by parallel computing multiple "heads", thereby enhancing the model's overall understanding ability of the text. The parallel computing feature of the multi-head attention mechanism helps to improve the processing speed of the model. By processing the information of multiple "heads" simultaneously, the model can complete complex text analysis tasks in a shorter time, thus improving the overall performance and enhancing the accuracy and reliability of the broadcast test.
[0102] In an embodiment of the present disclosure, multiple test information is processed based on the multi-head attention mechanism to obtain multiple broadcast information, including:
[0103] Feature extraction is performed on the test information to obtain multiple types of broadcast features;
[0104] Based on the multi-head attention mechanism, the user feature and multiple types of broadcast features are processed to obtain multiple first broadcast information;
[0105] The broadcast information is determined based on the multiple first broadcast information.
[0106] In an embodiment of the present disclosure, multiple types of broadcast features include content features, scenario features, and sender features;
[0107] Based on the multi-head attention mechanism, the user feature and multiple types of broadcast features are processed to obtain multiple first broadcast information, including:
[0108] Based on the multi-head attention mechanism, the content feature is processed to obtain content broadcast information;
[0109] Based on the multi-head attention mechanism, the user feature and the content feature are processed to obtain user broadcast information;
[0110] Based on the multi-head attention mechanism, the scenario feature and the content feature are processed to obtain scenario broadcast information;
[0111] Based on the multi-head attention mechanism, the sender feature and the content feature are processed to obtain sender broadcast information;
[0112] The content broadcast information, user broadcast information, scenario broadcast information, and sender broadcast information are all first broadcast information.
[0113] In this embodiment, multiple types of broadcast features may include: content features, scenario features, and sender features.
[0114] The content feature can be the text tone feature in the test content. The content feature can be extracted by methods such as keyword extraction and sentiment word analysis. According to the meaning expressed by the actual text information, the first broadcast information can be obtained. The first broadcast information can be a tone information obtained according to the content feature, that is, the content broadcast information.
[0115] The scene feature can be the application scene feature of the user. The scene feature can be extracted according to the scene keyword. According to the scene feature and content feature appearing in the text, the tone of the sender can be inferred, which is the scene broadcast information. For example, for a sentence "Have a meeting in the office at 3 pm", the corresponding tone should be serious and solemn. Different scenes can have corresponding different tones.
[0116] For example, for the sentence "I am in the amusement park", when not considering the scene feature and content feature, the output may be a relatively flat tone, but combined with the scene, it can be expressed as a happy and upward tone, which is more in line with the actual tone of the sender.
[0117] The sender feature can be the relationship or intimacy between the sender and the receiver. For example, if the receiver's note for the sender is "Mom", it can be determined as a family relationship at this time. When there are "eating" and "temperature drop" in the content feature, the tone can be defined as "concerned". When the receiver's note for the sender is "boss", it can be determined as a superior relationship at this time. When there are "documents" and "reports" in the content feature, the tone can be defined as "serious". According to the sender feature, a first broadcast information can be obtained. The first broadcast information can be a tone information, that is, the sender broadcast information.
[0118] The user feature can be the user's state. When the user is in a busy state and there are "as soon as possible" and "right away" in the content feature, the broadcast tone should be concise, clear, efficient and direct. When the user is relaxed, and there are "eating" and "happy" in the content feature, the broadcast tone can be relaxed, pleasant and amiable, that is, the user broadcast information.
[0119] The user feature is not obtained from the test information and does not belong to multiple types of features. In the test environment of this application, the user feature can be set in advance. In the application environment, for example, when applied in smart glasses, the user's state can be judged by monitoring the user's behavior, habits, etc.
[0120] Similarly, the above-mentioned sender feature is also set in advance, but not directly set, but indirectly set by setting notes and other methods, and then the sender feature is determined by the above-mentioned feature extraction and other methods.
[0121] In this embodiment, the multi-head attention mechanism can respectively judge and output the tones of user features, content features, scene features, and sender features, and the output results are multiple first broadcast messages.
[0122] It should be noted that in this embodiment, the content feature not only determines the first broadcast message (tone) corresponding to itself, but also participates in the determination process of the first broadcast messages of the other three features.
[0123] From the above, it can be concluded that the present disclosure extracts multiple types of broadcast features in the test information through feature extraction, such as content features, scene features, and sender features, etc., which together constitute the basis for generating broadcast messages. By considering user features and sender features in this embodiment, the method can generate more personalized broadcast messages that conform to the user's identity and relationship. The application of the multi-head attention mechanism enables the model to process multiple features in parallel, improving the processing speed and efficiency.
[0124] In an embodiment of the present disclosure, determining the broadcast message based on multiple first broadcast messages includes:
[0125] Determine the weight of the content broadcast message;
[0126] Respectively determine multiple degrees of association between user features, scene features, and sender features and content features;
[0127] Based on multiple degrees of association, determine the weights of user broadcast messages, scene broadcast messages, and sender broadcast messages;
[0128] Perform weighted calculation on multiple first broadcast messages to obtain the broadcast message.
[0129] In this embodiment, considering that the content feature is an important factor affecting the tone, and other factors can be understood as a supplement and correction to the content feature, the weight of the content broadcast message can be defined as , and the weight of the user broadcast message is defined as , the weight of the scene broadcast message is defined as , and the weight of the sender broadcast message is defined as .
[0130] In this embodiment, the weight of the content broadcast message is set as the first weight;
[0131] The sum of the weights of user broadcast messages, scene broadcast messages, and sender broadcast messages is set as the second weight;
[0132] The sum of the first weight and the second weight is 1.
[0133] In this embodiment, the first weight is , the first weight can be determined based on experiments.
[0134] The weights of user broadcast information, scenario broadcast information, and sender broadcast information can be allocated by calculating the correlation degrees between user features, scenario features, sender features, and content features.
[0135] The correlation degree can be calculated based on the feature matching degree or the Euclidean distance method, etc. For example, the content features include: "amusement park", "office", "cinema", "home", "mom", "sports". At this time, multiple scenarios appear, so the correlation degree between the scenario features and the content features is relatively high, and the correlation degrees between the user features and the sender features and the content features are relatively low.
[0136] Taking the calculation of the correlation degree between scenario features and content features as an example, convert the scenario features into vectors , convert the content features into vectors , calculate the distance according to the Euclidean distance calculation formula. If the dimensions (i.e., the number of elements) of the two vectors are different, they can be made to have the same dimension by filling in data, and then the Euclidean distance formula can be used for calculation. Usually, elements can be added to the vector with fewer dimensions, and the added element values can be filled in with default values or means.
[0137] The calculated Euclidean distance value can be used as a criterion for evaluating the correlation degree, and the correlation degree can be obtained by presetting a mapping table between the Euclidean distance value and the correlation degree.
[0138] The correlation degrees between user features, scenario features, sender features, and content features can be defined as , , , then the weight corresponding to the user feature is , the weight corresponding to the scenario feature is , the weight corresponding to the sender feature is .
[0139] For example, the calculated correlation degree between user features and content features is 0.2, the correlation degree between scenario features and content features is 0.8, the correlation degree between sender features and content features is 0.2, and the first weight is defined as 0.4. At this time, the total weights of user features, scenario features, and sender features can be 0.6.
[0140] Then the weight corresponding to the user feature is 0.1, the weight corresponding to the scenario feature is 0.4, the weight corresponding to the sender feature is 0.1, and the weight corresponding to the content feature is 0.4 (i.e., ).
[0141] The weighted sum of the user broadcast information, scene broadcast information, sender broadcast information, and content broadcast information can be calculated according to the calculated weights to obtain the final target broadcast information.
[0142] As can be seen from the above, the present disclosure determines the weight of the content broadcast information and assigns the weights of the user broadcast information, scene broadcast information, and sender broadcast information based on the correlation between other features and content features, so as to more accurately reflect the tone and intention in the input text, make the broadcast information more in line with the actual context, and reduce the possibility of misunderstanding and ambiguity. By considering the correlation between user features, scene features, sender features, and content features, the method of the present disclosure can generate more personalized broadcast information that conforms to the user's identity and relationship. The weight allocation mechanism of the present disclosure enables the model to pay more attention to key information when processing multiple broadcast information, thereby improving the processing efficiency and accuracy, and enhancing the accuracy and reliability of the broadcast test.
[0143] Corresponding to the broadcast test method in the above embodiment, Figure 2 It is a structural block diagram of a broadcast test device provided by an embodiment of the present disclosure. For the convenience of description, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The broadcast test device 20 includes: a model construction module 21, a model test module 22, and a model adjustment module 23.
[0144] Among them, the model construction module 21 is used to construct a broadcast test model;
[0145] The model test module 22 is used to input multiple test information into the broadcast test model to determine multiple target broadcast information. The multiple test information is the information in the test set, and the test set includes multiple test information and the corresponding broadcast information for each test information;
[0146] The model adjustment module 23 is used to, in response to the number of error information in the multiple target broadcast information being greater than a first preset value and the type of the error information being a format error, adjust the broadcast test model based on a first method; and execute the step of inputting the multiple test information into the broadcast test model to determine the multiple target broadcast information until the number of error information in the multiple target broadcast information is less than or equal to the first preset value;
[0147] In response to the number of error information in the multiple target broadcast information being greater than a first preset value and the type of the error information being a semantic error, adjust the broadcast test model based on a second method; and execute the step of inputting the multiple test information into the broadcast test model to determine the multiple target broadcast information until the number of error information in the multiple target broadcast information is less than or equal to the first preset value;
[0148] Among them, the processing methods of the first method and the second method are different.
[0149] In one embodiment of the present disclosure, the broadcast test model includes a broadcast feature library;
[0150] The model adjustment module 23 is specifically configured to process error information based on a regular expression and adjust the broadcast feature library based on the processed error information;
[0151] The model adjustment module 23 is specifically further configured to determine semantic features in the error information and adjust the broadcast feature library based on the semantic features.
[0152] In one embodiment of the present disclosure, the model adjustment module 23 is specifically further configured to, in response to a semantic error being an error in text semantic expression, extract semantic features of the text and add the semantic features to the broadcast feature library;
[0153] In response to the semantic error being a non - text semantic expression error, obtain the initial semantics and correct semantics of the non - text and add the initial semantics and correct semantics to the broadcast feature library.
[0154] In one embodiment of the present disclosure, the model test module 22 is specifically configured to input multiple test messages into the language processing module, process the multiple test messages based on the multi - head attention mechanism, and obtain multiple broadcast messages;
[0155] Process the multiple test messages based on the multi - head attention mechanism to obtain multiple broadcast messages;
[0156] The language processing module is a module in the broadcast test model; the language processing module is constructed based on the multi - head attention mechanism.
[0157] In one embodiment of the present disclosure, the model test module 22 is specifically further configured to perform feature extraction on the test messages to obtain multiple types of broadcast features;
[0158] Process the user features and multiple types of broadcast features based on the multi - head attention mechanism to obtain multiple first broadcast messages;
[0159] Determine the broadcast message based on the multiple first broadcast messages.
[0160] In one embodiment of the present disclosure, the multiple types of broadcast features include content features, scenario features, and sender features;
[0161] The model test module 22 is specifically further configured to process the content features based on the multi - head attention mechanism to obtain content broadcast messages;
[0162] Process the user features and content features based on the multi - head attention mechanism to obtain user broadcast messages;
[0163] Process the scenario features and content features based on the multi - head attention mechanism to obtain scenario broadcast messages;
[0164] Process the sender features and content features based on the multi-head attention mechanism to obtain the sender's broadcast information;
[0165] The content broadcast information, user broadcast information, scenario broadcast information, and sender broadcast information are all the first broadcast information.
[0166] In an embodiment of the present disclosure, the model testing module 22 is specifically further configured to determine the weight of the content broadcast information;
[0167] Respectively determine multiple degrees of association between the user features, scenario features, and sender features and the content features;
[0168] Determine the weights of the user broadcast information, scenario broadcast information, and sender broadcast information based on the multiple degrees of association;
[0169] Perform weighted calculation on the multiple first broadcast information to obtain the broadcast information.
[0170] See Figure 3 , Figure 3 which is a schematic block diagram of a wearable device provided in an embodiment of the present disclosure. As Figure 3 shown, the wearable device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through the communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 the functions of the modules 21 to 23 shown.
[0171] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0172] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the orientation information of the fingerprint of a user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0173] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may further include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0174] In a specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the broadcast test method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the wearable device described in the embodiments of the present disclosure, which will not be elaborated herein.
[0175] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It may also be completed by instructing relevant hardware through the computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0176] A computer-readable storage medium may be an internal storage unit of the wearable device in any of the foregoing embodiments, such as the hard disk or memory of the wearable device. The computer-readable storage medium may also be an external storage device of the wearable device, such as a plug-in hard disk equipped on the wearable device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the wearable device. The computer-readable storage medium is used to store computer programs and other programs and data required by the wearable device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0177] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.
[0178] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working processes of the wearable device and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0179] In several embodiments provided in this application, it should be understood that the disclosed wearable device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, and can also be electrical, mechanical, or other forms of connection.
[0180] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this disclosure.
[0181] In addition, in each embodiment of the present disclosure, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0182] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A broadcast test method, characterized in that, Including: Construct a broadcast test model; Input multiple test messages into the broadcast test model to determine multiple target broadcast messages. The multiple test messages are messages in a test set, and the test set includes multiple test messages and the corresponding broadcast messages for each test message; In response to the number of error messages in the multiple target broadcast messages being greater than a first preset value and the type of the error message being a format error, adjust the broadcast test model based on a first method; And, execute the step of inputting multiple test messages into the broadcast test model to determine multiple target broadcast messages until the number of error messages in the multiple target broadcast messages is less than or equal to the first preset value; In response to the number of error messages in the multiple target broadcast messages being greater than a first preset value and the type of the error message being a semantic error, adjust the broadcast test model based on a second method; And, execute the step of inputting multiple test messages into the broadcast test model to determine multiple target broadcast messages until the number of error messages in the multiple target broadcast messages is less than or equal to the first preset value; Wherein, the processing methods of the first method and the second method are different; The step of inputting multiple test messages into the broadcast test model to determine multiple target broadcast messages includes: Input multiple test messages into a language processing module, and process the multiple test messages based on a multi-head attention mechanism to obtain multiple target broadcast messages; Extract feature extractions from the test messages to obtain content features, scenario features, and sender features; Process the user features and content features based on a multi-head attention mechanism to obtain user broadcast messages; Process the scenario features and content features based on a multi-head attention mechanism to obtain scenario broadcast messages; Process the sender features and content features based on a multi-head attention mechanism to obtain sender broadcast messages; The user broadcast messages, scenario broadcast messages, and sender broadcast messages are all first broadcast messages; Determine the broadcast message based on multiple first broadcast messages; The language processing module is a module in the broadcast test model; the language processing module is constructed based on a multi-head attention mechanism.
2. The broadcast test method according to claim 1, characterized in that The broadcast test model includes a broadcast feature library; The step of adjusting the broadcast test model based on the first method includes: Process the error message based on a regular expression, and adjust the broadcast feature library based on the processed error message; The step of adjusting the broadcast test model based on the second method includes: Determine the semantic features in the error message, and adjust the broadcast feature library based on the semantic features.
3. The broadcast test method according to claim 2, characterized in that, The step of determining the semantic features in the error message and adjusting the broadcast feature library based on the semantic features includes: In response to the semantic error being a text semantic expression error, extract the semantic features of the text and add the semantic features to the broadcast feature library; In response to the semantic error being a non-text semantic expression error, obtain the initial semantics and correct semantics of the non-text, and add the initial semantics and the correct semantics to the broadcast feature library.
4. The broadcast test method according to claim 1, wherein The step of determining the broadcast message based on multiple first broadcast messages includes: Determine the weight of the content broadcast message; the content broadcast message belongs to the first broadcast message; Determine multiple degrees of association between the user characteristics, the scenario characteristics, and the sender characteristics and the content characteristics respectively; Determine the weight of the user broadcast information, the weight of the scenario broadcast information, and the weight of the sender broadcast information based on the multiple degrees of association; Perform weighted calculation on the multiple first broadcast information to obtain the broadcast information.
5. A broadcast test device, characterized in that, Including: A model construction module for constructing a broadcast test model; A model test module for inputting multiple test information into the broadcast test model to determine multiple target broadcast information, where the multiple test information is information in a test set, and the test set includes multiple test information and the corresponding broadcast information for each test information; A model adjustment module for, in response to the number of error information in the multiple target broadcast information being greater than a first preset value and the type of the error information being a format error, adjusting the broadcast test model based on a first method; And, execute the step of inputting multiple test information into the broadcast test model to determine multiple target broadcast information until the number of error information in the multiple target broadcast information is less than or equal to the first preset value; In response to the number of error information in the multiple target broadcast information being greater than a first preset value and the type of the error information being a semantic error, adjust the broadcast test model based on a second method; And, execute the step of inputting multiple test information into the broadcast test model to determine multiple target broadcast information until the number of error information in the multiple target broadcast information is less than or equal to the first preset value; Wherein, the processing methods of the first method and the second method are different; The model test module is specifically configured to input multiple test information into a language processing module, and process the multiple test information based on a multi-head attention mechanism to obtain multiple target broadcast information; Extract features from the test information to obtain content features, scenario features, and sender features; Process the user characteristics and the content characteristics based on a multi-head attention mechanism to obtain user broadcast information; Process the scenario characteristics and the content characteristics based on a multi-head attention mechanism to obtain scenario broadcast information; Process the sender characteristics and the content characteristics based on a multi-head attention mechanism to obtain sender broadcast information; The user broadcast information, the scenario broadcast information, and the sender broadcast information are all first broadcast information; Determine the broadcast information based on the multiple first broadcast information; The language processing module is a module in the broadcast test model; the language processing module is constructed based on a multi-head attention mechanism.
6. A wearable device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Assessing the subjective quality of TTS systems which accounts for variations between synthesised and original speech
GB2423903A