Content matching method and apparatus, electronic device, and computer readable medium
By acquiring and matching the set of voice content and the set of target description information in the voice interaction device, and taking into account the mixing and multiple sounds of Pinyin, the problem of inaccurate voice recognition is solved, and the effectiveness and experience of user operation are improved.
Patent Information
- Application Number
- CN202111296077.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2041-11-03
AI Technical Summary
Inaccurate matching between voice recognition results and information displayed on interactive devices can prevent users from effectively executing their intended actions, thus impacting the user experience.
By acquiring the set of voice content input by the user and the set of target description information, a matching operation is performed to determine the matching result for each content, including identical, equivalent, and non-matching. The matching degree is determined based on the matching result, taking into account the mixing and multiple pronunciations of Pinyin to improve matching accuracy.
It improves the accuracy of matching the speech recognition content with the target description information, ensuring that the user's intended actions can be effectively executed.
Smart Images

Figure CN114093355B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a content matching method, apparatus, electronic device, and computer-readable medium. Background Technology
[0002] With the development of various new forms of interaction, voice interaction remains the mainstream form of interaction. In interaction scenarios, users often use voice to describe specific information on the interface of the interactive device to guide the device to perform corresponding operations. However, the current voice recognition results and the information displayed by the interactive device are not accurately matched, which makes it impossible to effectively execute the user's intention operation and affects the user experience. Summary of the Invention
[0003] This application proposes a content matching method, apparatus, electronic device, and computer-readable medium to improve the above-mentioned deficiencies.
[0004] In a first aspect, embodiments of this application provide a content matching method, comprising: acquiring a first content set, the first content set being acquired in advance based on user-input speech, the first content set including at least one first content, the at least one first content including the first pinyin of a first Chinese character within the speech; acquiring a second content set, the second content set including at least one second content of target description information, the at least one second content including the second pinyin of a second Chinese character corresponding to the target description information, the target description information corresponding to a target operation; performing a matching operation for each first content and the second content set, determining a matching result for each first content, the matching result including identical, equivalent matching, and non-matching, the equivalent matching including at least one of mixed-sound matching and multi-sound matching; determining the matching degree of the first content based on the matching result of the first content, different matching results having different matching degrees; determining the total matching degree of the first content set based on the matching degree of each first content, the total matching degree being used as the matching result between the first content set and the second content set.
[0005] Secondly, embodiments of this application also provide a content matching apparatus, including: a first acquisition unit, a second acquisition unit, a determination unit, a statistics unit, and a calculation unit. The first acquisition unit is used to acquire a first content set, which is pre-acquired based on user-inputted speech. The first content set includes at least one first content, and the at least one first content includes the first pinyin of a first Chinese character within the speech. The second acquisition unit is used to acquire a second content set, which includes at least one second content of target description information. The at least one second content includes the second pinyin of a second Chinese character corresponding to the target description information, and the target description information corresponds to a target operation. The determination unit is used to perform a matching operation between each first content and the second content set, determining a matching result for each first content. The matching result includes identical, equivalent matching, and non-matching, and the equivalent matching includes at least one of mixed-sound matching and multi-sound matching. The statistics unit is used to determine the matching degree of the first content based on the matching result of the first content, where different matching results have different matching degrees. The calculation unit is used to determine the total matching degree of the first content set based on the matching degree of each first content, and the total matching degree is used as the matching result between the first content set and the second content set.
[0006] Thirdly, embodiments of this application also provide an electronic device, including: one or more processors; a memory; one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to perform the methods described above.
[0007] Fourthly, embodiments of this application also provide a computer-readable medium storing processor-executable program code, which, when executed by the processor, causes the processor to perform the above-described method.
[0008] Fifthly, embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the above-described method.
[0009] The content matching method, apparatus, electronic device, computer-readable medium, and computer program product provided in this application perform a matching operation on each first content set and the second content set to determine the matching result of each first content set. Since the matching result includes identical, equivalent, and non-matching, and the equivalent matching includes at least one of mixed-sound matching and multi-sound matching, compared to simply determining identical or different matching results, it can take into account the effects of mixed-sound and multi-sound matching of pinyin, so that the matching result also includes at least one of mixed-sound matching and multi-sound matching. Therefore, when determining the matching degree of each first content set based on the matching result and then determining the total matching degree of the first content set, the determination of the total matching degree can be more accurate, thereby improving the accuracy of matching the relationship between speech recognition content and target description information. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A schematic diagram of an interactive scenario provided by an embodiment of this application is shown;
[0012] Figure 2 A flowchart of a content matching method provided in an embodiment of this application is shown;
[0013] Figure 3 A flowchart of a content matching method provided in another embodiment of this application is shown;
[0014] Figure 4 A flowchart of a content matching method provided in another embodiment of this application is shown;
[0015] Figure 5 An embodiment of this application is shown. Figure 4 Implementation method of S404 in the document;
[0016] Figure 6 Another embodiment of this application is shown. Figure 4 Implementation method of S404 in the document;
[0017] Figure 7 This application illustrates another embodiment of the present application. Figure 4 Implementation method of S404 in the document;
[0018] Figure 8A flowchart of a content matching method provided in another embodiment of this application is shown;
[0019] Figure 9 A block diagram of a content matching apparatus provided in one embodiment of this application is shown;
[0020] Figure 10 A structural block diagram of an electronic device provided in an embodiment of this application is shown;
[0021] Figure 11 This is a storage unit in this application embodiment for storing or carrying program code that implements the content matching method according to this application embodiment. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0023] With the development of various new interaction methods, voice interaction remains the mainstream interaction method. Interaction scenarios often involve users using voice to describe specific information on the device's interface to guide the device in performing corresponding operations. Please refer to [reference needed]. Figure 1 , Figure 1 An interactive scenario is illustrated, in which a user faces an operation interface 101 displayed on an electronic device. The operation interface 101 displays multiple controls and content that can be used to perform specific operations. For example, if a control is labeled "Settings," clicking it will take the user to the settings interface. To facilitate voice control of the operation interface 101 and the corresponding operations on the electronic device, the user can input voice commands which are then recognized by the electronic device. The electronic device can then pre-obtain descriptive information corresponding to the operation interface. This descriptive information can be the descriptive information of the controls on the interface, such as the displayed content of the controls, or the descriptive content of the gestures, such as "double-tap" or "swipe left."
[0024] Therefore, when an electronic device recognizes the user's voice input and matches the voice content with one of the multiple descriptive information items on the user interface, it can execute the specified operation corresponding to that descriptive information. For example, if the user inputs "viewing history," the electronic device recognizes that the voice content matches the "viewing history" content on the user interface, and then executes the operation corresponding to the control for that viewing history, i.e., opens the viewing history interface.
[0025] Then, due to the user's inaccurate pronunciation, the speech recognition is inaccurate, which leads to a mismatch between the speech recognition result and the information displayed on the interactive device. As a result, the user's intended operation cannot be effectively executed, affecting the user experience.
[0026] Therefore, in order to overcome the above-mentioned defects, the embodiments of this application provide a content matching method, apparatus, electronic device, computer-readable medium, and computer program product, which can fully consider the matching problems in cases of polyphonic characters, easy confusion, and incomplete user descriptions, and improve the accuracy of matching the relationship between speech recognition content and target description information.
[0027] Please see Figure 2 , Figure 2 A content matching method is shown, applied to the aforementioned electronic device, the method comprising: S201 to S205.
[0028] S201: Get the first content set.
[0029] The first content set is obtained in advance based on the user's voice input, and the first content set includes at least one first content, wherein the at least one first content includes the first pinyin of the first Chinese character in the voice.
[0030] In one implementation, the electronic device displays an operating interface. For example, the electronic device may include a display screen on which the operating interface is displayed; or, for another example, the electronic device may be a projection device on which the operating interface is projected. The user inputs voice based on the operating interface, and this voice can be captured by the electronic device's audio acquisition device to obtain audio data. Alternatively, the user may wear an audio acquisition device that can use the user's voice input to obtain audio data and send the audio data to the electronic device.
[0031] Electronic devices, based on natural language processing technology, can convert audio data into a text sequence. This text sequence can include Chinese characters, letters, numbers, punctuation marks, and special characters, thus obtaining a speech-text sequence corresponding to the user's input speech content. The special characters can be predefined symbols, such as parentheses, dashes, and other non-punctuation and non-text symbols. Numbers, punctuation marks, and special characters may also be excluded from the text sequence; that is, the text sequence may only include Chinese characters and letters, meaning it may or may not include other characters besides Chinese characters.
[0032] In one implementation, the electronic device may have a built-in voice processing module. This module can perform natural language processing on the audio data to obtain the corresponding text sequence. Alternatively, the electronic device may send the audio data to another device, which will then process the audio data to obtain the corresponding text sequence and return it to the electronic device.
[0033] As one implementation method, the method for determining the speech-text sequence is to identify the user's speech to obtain a first initial text sequence. The first initial text sequence includes Chinese characters, letters, numbers, punctuation marks, and other special characters. As another implementation method, the special characters and numbers in the first initial text sequence can be removed, leaving only Chinese characters, English characters, and punctuation marks as the speech-text sequence. The Chinese characters in the speech-text sequence are then converted into Pinyin, while the others remain unchanged, to obtain the first content set.
[0034] S202: Obtain the second content set.
[0035] The electronic device can display content within the device's operating interface and associate the displayed content with the corresponding operation to obtain a text sequence for the operating interface. This text sequence can include Chinese characters, letters, numbers, punctuation marks, and special characters corresponding to the displayed content within the operating interface. Similarly, special characters and numbers can be removed from the text sequence, leaving only Chinese characters, English characters, and punctuation marks as a new text sequence. The Chinese characters in this new text sequence are then converted to Pinyin, while other elements remain unchanged, to obtain a second content set.
[0036] As one implementation method, the user interface can be developed by setting corresponding description information for each control's operation or other operations on the interface. Alternatively, the electronic device can scan the content displayed on the user interface to obtain the description information for each control and associate each description information with the operation of that control. In some embodiments, when the user interface is displayed, at least a portion of the description information from multiple description information sets will be displayed, for example, Figure 1 The interface displays descriptive information for each space.
[0037] In one implementation, the target description information is one of a plurality of description information. The electronic device will match the first content set with the second content set corresponding to each description information to find the description information that successfully matches the first content set. The target description information is the description information of the current operation.
[0038] S203: Perform a matching operation on each of the first content and the second content set to determine the matching result for each of the first content, wherein the matching result includes identical, equivalent match, and no match.
[0039] The equivalent matching includes at least one of mixed-tone matching and multi-tone matching.
[0040] As one implementation method, since both the first content set and the second content set can include the pinyin corresponding to Chinese characters, characters of a specified type, punctuation marks, etc., where the specified type refers to characters that are not Chinese characters, when matching the first content in the first content set and the second content in the second content set, there will be mutual matching between the pinyin corresponding to Chinese characters, characters of a specified type, punctuation marks, etc. The corresponding matching results will include identical, equivalent matching, and non-matching. The specific matching process can be referred to in subsequent embodiments.
[0041] In some embodiments, a matching result of "the two are the same" means that the first content and the second content are completely identical. For example, if the first content is the pinyin "shang" and the second content is also the pinyin "shang", then the matching result of the first content and the second content is "the two are the same". A mismatch means that the matching result is neither the two are the same nor an equivalent match. An equivalent match lies between "the two are the same" and "the two are not the same", indicating that there is some similarity or correlation between the two, but they are not completely identical nor completely unrelated. Specifically, an equivalent match includes at least one of mixed-sound matching and multi-sound matching.
[0042] As one implementation method, sound mixing matching is used to characterize a sound mixing relationship between the first content and the second content, wherein the sound mixing relationship includes at least one of front and back nasal sounds, nasal lateral sounds, and retroflex consonants. Specifically, as shown in Table 1:
[0043] Table 1
[0044] Types of mixing relationships Specific Pinyin Groups Front and back nasal sounds (eng, en), (ing, in), (ang, an), etc. Nasal consonants and lateral consonants (l, n) etc. flat and retroflex tongue (z, zh), (c, ch), (s, sh), (zh, ch, sh), etc.
[0045] Table 1 above lists three types of mixed sound relationships, namely front and back nasal sounds, nasal and lateral sounds, and flat and retroflex sounds, as well as at least partial pinyin groups corresponding to each type. Then, when the first content is the first pinyin and the second content is the second pinyin, if a matching pinyin group can be found for the first pinyin and the second pinyin, it can be determined that the first pinyin and the second pinyin belong to a mixed sound relationship, that is, the matching result of the two is a mixed sound match. In addition, it should be noted that the initial and final sounds of the first pinyin can be respectively matched with the initial and final sounds of the second pinyin to determine whether a pinyin group in Table 1 can be found. For example, if the first pinyin is yang and the second pinyin is yan, the initial and final sounds corresponding to the first pinyin are y and ang respectively, and the initial and final sounds corresponding to the second pinyin are y and an respectively. Then, it can be determined that ang of the first pinyin and an of the second pinyin can find the corresponding pinyin group (ang, an) in Table 1, and they belong to the front and back nasal sound relationship. Therefore, it can be determined that the matching result of the first pinyin and the second pinyin is a mixed sound match. Furthermore, it should be noted that the flat and retroflex sound relationship can include flat tongue and retroflex tongue, as well as both being retroflex tongues. That is, one of the first pinyin and the second pinyin contains a flat tongue and the other contains a retroflex tongue, or both the first pinyin and the second pinyin contain retroflex tongues, and both belong to the flat and retroflex sound relationship. Moreover, only some of the pinyin combinations in Table 1 above are shown, and more combinations can be added according to requirements during actual use, which is not limited here.
[0046] As an implementation manner, the multiple-pinyin match is used to represent that at least one of the first Chinese character corresponding to the first content and the second Chinese character corresponding to the second content corresponds to multiple pinyins, and all the pinyins corresponding to the first Chinese character are the first pinyin, all the pinyins corresponding to the second Chinese character are the second pinyin, and there are identical pinyins among the first pinyin and the second pinyin.
[0047] Specifically, at least one of the first Chinese character corresponding to the first content and the second Chinese character corresponding to the second content is a polyphonic character of a target type. A polyphonic character of a target type means that the pronunciation of the Chinese character has multiple variations, and there are also multiple pinyins corresponding to the polyphonic character. For example, the pinyins corresponding to the pronunciation of "重" include chong and zhong, so "重" is a polyphonic character of a target type. Therefore, "重" corresponds to multiple pinyins, and "multiple" here means at least two. So, if the first Chinese character is "重", the first pinyin corresponding to the first Chinese character is chong and zhong. When matching, both chong and zhong need to be matched with the second pinyin. As long as there is at least one identical pinyin between all the pinyins corresponding to the first pinyin and all the pinyins corresponding to the second pinyin, it is considered that the matching result of the two is a multiple-pinyin match.
[0048] In addition, for text and punctuation marks of a specified type, the matching result is either the two are the same or they are not matched. If the first content and the second content are both text of the specified type, the matching result is the same when they are completely identical, otherwise they are not matched. If the first content and the second content are both punctuation marks, the matching result is considered to be the same, that is, if they are both of the punctuation mark type, they are considered to be the same regardless of whether their punctuation marks are the same.
[0049] Specifically, the matching process of the first and second contents described above can be referred to in subsequent embodiments, and will not be described here.
[0050] As one implementation method, the specific implementation method for performing the matching operation between each first content and the second content set is as follows: select a first content from the first content set as the first target content; match the first target content with the second content set to obtain the matching result of the first target content; based on a preset matching order, select the next first content from the first content set as the new first target content, so as to traverse each first content in the first content set and determine the matching result of each first content.
[0051] Specifically, it can be that the first content is matched with every second content in the second content set to obtain a matching result between the first content and each second content; or it can be that the first content is matched with each second content in the second content set according to a preset matching order. When a successful match is achieved, a specified position of the second content corresponding to the successful match is determined. Second content after the specified position in the preset matching order is no longer matched. Here, a successful match includes identical or equivalent matches, i.e., no non-match. As one implementation method, when determining the aforementioned speech-text sequence,
[0052] As an implementation, a first serial number is set for each character in the speech text sequence. Here, the character can be a Chinese character, a specified type of text, a punctuation mark, etc. Then, the Chinese characters in the speech text sequence are converted into pinyin to obtain a first content set. Each first content also corresponds to a serial number. In some embodiments, the serial numbers of the respective first contents in the first content set correspond to the first sequence of the respective characters in the speech text sequence. For example, if the user inputs the speech "I love Chongqing CQ", the recognized speech text sequence is {"我", "爱", ",", "重", "庆", "C", "Q"}. The Chinese characters in the speech text sequence are converted into pinyin, and the obtained first content set is {"wo", "ai", ",", "zhong", "qing", "C", "Q"}. In some embodiments, consecutive specified types of text can be grouped together and correspond to the same serial number. For example, CQ can be grouped together, that is, the speech text sequence is {"我", "爱", ",", "重", "庆", "CQ"}, and the first content set is {"wo", "ai", ",", "zhong", "qing", "CQ"}. In some other embodiments, the Chinese characters in the speech text sequence may include polyphonic characters of the above target type. Then, all the pinyin corresponding to the multiple characters of the target type are used as the pinyin of the Chinese character, and all the pinyin correspond to the same serial number. For example, the "重" in the speech text sequence is a polyphonic character of the target type, and the pinyin corresponding to the "重" is chong and zhog. Then, the first content set can be {"wo", "ai", ",", "[zhong,chong]", "qing", "CQ"}, and the serial numbers of the first content set are 1, 2, 3, 4, 5, 6 in sequence, corresponding to the 6 elements of {"wo", "ai", ",", "[zhong,chong]", "qing", "CQ"} respectively. Thus, each first content in the first content set corresponds to a serial number.
[0053] Similarly, the serial numbers of the respective second contents in the second content set can also be set in the above manner. Then, the serial numbers of the respective first contents in the first content set are recorded as the first serial numbers, and the serial numbers of the respective second contents in the second content set are recorded as the second serial numbers. The preset matching order can be in ascending order of the serial numbers. Of course, it can also be in descending order, which is not limited here.
[0054] After obtaining the matching result of each first content in the above manner, S204 is executed to obtain the matching degree of each first content.
[0055] S204: Determine the matching degree of the first content based on the matching result of the first content.
[0056] It should be noted that the matching degree of different matching results can be different. In some embodiments, the matching degree corresponding to identical, equivalent, and non-matching results decreases in that order. Specifically, the matching degree corresponding to polyphonic matching is greater than that of mixed-tone matching. As one implementation method, different scores can be set for different matching results, and this score serves as the matching degree of the matching result. The higher the score, the higher the matching degree, which indicates a higher similarity or relevance between the first and second content. For example, the matching degree corresponding to identical results is 1, the matching degree corresponding to non-matching results is 0, and the matching degree corresponding to equivalent matching results is between 0 and 1. For instance, the matching degree corresponding to polyphonic matching results is 0.8, and the matching degree corresponding to mixed-tone matching results is 0.4. Of course, the matching degree corresponding to these different matching results can also be set to other values while satisfying the above-mentioned size relationship; this is not limited here.
[0057] In one implementation, a matching operation is performed on each of the first content items and the second content set to determine the matching result of each of the first content items. The matching degree of the first content item is determined based on the matching result of the first content item by selecting a first content item from the first content set as the first target content item, matching the first target content item with the second content set to obtain the matching result of the first target content item, determining the matching degree of the first target content item based on the matching result of the first target content item, and selecting the next first content item from the first content set as the new first target content item based on a preset matching order, so as to traverse each of the first content items in the first content set and determine the matching degree of each of the first content items.
[0058] In one implementation, the first target content is matched with the second content set in a first manner. That is, after the first target content is successfully matched with a second content in the second content set, the remaining unmatched second content is no longer matched. The matching result corresponding to the successful match is taken as the matching result of the first target content. For example, if the first target content is successfully matched with the second second content in the second content set and the matching result is a mixed match, then the second target content does not need to be matched with the second content after the second second content. The matching result of the second target content is a mixed match and the corresponding matching degree is the matching degree of the mixed match, for example, 0.4. Then the matching degree of the first target content can be recorded as 0.4. Then, the matching operation is performed on the next first content after the first target content according to the preset matching order.
[0059] As another implementation method, the first target content is matched with the second content set in a second way, that is, the first target content is matched with each second content in the second content set, and then the matching result of the first target content and each second content is obtained, and the matching degree of each matching result is accumulated as the matching degree of the first target content.
[0060] In the embodiments of this application, each embodiment is described in a first manner by matching the first target content with the second content set. However, it should be understood that the embodiments of this application do not limit the matching of the first target content with the second content set to the first manner.
[0061] S205: Determine the total matching degree of the first content set based on the matching degree of each of the first contents.
[0062] As one implementation method, the matching degree of each first content in the first content set can be summed to obtain the total matching degree of the first content set.
[0063] As another implementation, to avoid differences and inaccuracies in matching scores due to varying speech lengths, a first quantity of the first content within the first content set is obtained. Specifically, this first quantity can be the largest sequence number among the aforementioned first sequences, i.e., the first quantity is the total number of characters in the speech-text sequence. For example, if the first content set is {“wo”, “ai”, “,”, “[zhong,chong]”, “qing”, “CQ”}, the first quantity of the first content set is 6. Then, the matching scores of each of the first contents are summed to obtain the total matching score. Based on the total matching score and the first quantity, the total matching score is obtained. As one implementation, the ratio of the total matching score to the first quantity can be used as the total matching score, thereby normalizing the total matching score and avoiding excessively high matching scores due to excessively long speech content.
[0064] As another implementation method, the length of the interface text sequence, i.e., the length of the second content set, also needs to be considered. Then, the second quantity of the second content within the second content set is obtained. Specifically, this second quantity can be the largest sequence number among the aforementioned second sequences, i.e., the second quantity is the total number of characters in the interface text sequence. Then, a first matching degree is obtained based on the total matching degree and the first quantity, and a second matching degree is obtained based on the total matching degree and the second quantity. The total matching degree is then obtained based on the first and second matching degrees. As another implementation method, the ratio of the total matching degree to the first quantity can be used as the first matching degree, and the ratio of the total matching degree to the second quantity can be used as the second matching degree. The first and second matching degrees are then weighted and summed to obtain the total matching degree. Therefore, the influence of both the speech text sequence and the interface text sequence on the total matching degree is considered, avoiding excessively long speech text sequences or interface text sequences that could lead to an excessively large cumulative total matching degree due to length limitations, thus further optimizing the recognition results.
[0065] It is understood that the total matching degree is the matching result between the first content set and the second content set. When there are multiple candidate description information, one description information is selected from the multiple candidate description information as the target description information. The operation corresponding to the target description information is the target operation. The content matching method of this application embodiment is executed for the target description information. Then, the matching result between the first content set and the second content set, i.e., the total matching degree of the first content set, is obtained. If the total matching degree meets the preset conditions, the target operation is executed. If the preset conditions are not met, the next description information is used as the new target description information, and the execution of the content matching method of this application embodiment is returned until a total matching degree that meets the preset conditions is found or all description information is matched. By obtaining the matching results between the first content set and each descriptive information, the descriptive information corresponding to the total matching degree that meets the preset conditions can be found. Then, the operation corresponding to the descriptive information is executed. The total matching degree meeting the preset conditions can be determined as the equivalence between the first content set and the second content set corresponding to the target descriptive information, that is, the user-input voice content is equivalent to the target descriptive information. It should be noted that the equivalence between the first content set and the second content set corresponding to the target descriptive information can mean that the user-input voice content includes the target descriptive information. The specific implementation method for the total matching degree meeting the preset conditions is to determine the total matching degree between the first content set and each descriptive information, and take the descriptive information corresponding to the highest total matching degree as the target descriptive information. Thus, the descriptive information that is equivalent to the user-input voice content among all the descriptive information can be determined, and the operation corresponding to the descriptive information is then executed.
[0066] Therefore, in the method provided by the embodiments of the present application, since the matching result includes that they are the same, equivalent matching, and non-matching, and the equivalent matching includes at least one of mixed sound matching and multi-tone matching, compared with simply determining the matching result of the same or different, the effects of mixed sound and multi-tone of pinyin can be considered, so that the matching result also includes at least one of mixed sound matching and multi-tone matching. Therefore, when determining the matching degree of each first content based on the matching result and then determining the total matching degree of the first content set, the determination of the total matching degree can be more accurate, and the accuracy of the matching relationship between the speech recognition content and the target description information can be improved.
[0067] Please refer to Figure 3 , Figure 3 which shows a content matching method provided by the embodiments of the present application and applied to the above electronic device. The method includes: S301 to S313.
[0068] S301: Obtain a first content set.
[0069] S302: Obtain a second content set.
[0070] S303: Select a first content from the first content set as the first target content.
[0071] As an implementation manner, the above-mentioned manner can be adopted to set a first serial number for each first content in the first content set. For example, if the first content set is {"wo", "ai", ",", "[zhong,chong]", "qing", "CQ"}, the first serial numbers of each first content in the first content set are respectively "wo" is 1, "ai" is 2, "," is 3, "[zhong,chong]" is 4, "qing" is 5, "CQ" is 6. That is to say, the serial numbers corresponding to each character in the speech text sequence {"我", "爱", ",", "重", "庆", "CQ"} corresponding to this first content set are also 1, 2, 3, 4, 5, 6 in sequence.
[0072] It can be understood that a preset matching order can be set in advance, that is, each first content in the first content set is matched with the second content set one by one according to this preset matching order. Exemplarily, this preset matching order can be the order of gradually increasing from small to large according to the first serial number, that is, "wo", "ai", ",", "[zhong,chong]", "qing", "CQ" are used as the first target content to be matched with the second content set in sequence.
[0073] S304: Match the first target content with the second content set to obtain the matching result of the first target content.
[0074] Similarly, each second content item in the second content set can also have its second number set using the same first numbering method as described above. In this case, the preset matching order of the second content set can be in the order of increasing second numbers from small to large.
[0075] As one implementation method, a second content is selected from the second content set as the second target content, and the first target content is matched with the second target content. In the embodiments of this application, the first method described above can be used to perform the matching of the first target content with the second content set.
[0076] It should be noted that the first target content set is the first content in the first content set that needs to be matched, that is, the first content of the current operation, and the second target content is the second content that matches the first target content, that is, the second content of the current operation.
[0077] S305: Based on the matching result of the first target content, determine the initial matching degree of the first target content.
[0078] The process of determining the initial matching degree of the first target content can refer to the process mentioned above, which describes different matching results corresponding to different matching degrees and the process of determining the matching degree based on the matching results, namely, the steps in S204 mentioned above.
[0079] S306: Determine whether the matching result of the first target content is a successful match.
[0080] Understandably, a successful match can be defined as a match where the two are identical or equivalent. Specifically, if the match between the first target content and at least one of the second content in the second content set is identical, a mixed-tone match, or a multi-tone match, then the match result of the first target content is considered a successful match. If the match result of the first target content is a successful match, then S307 is executed; if the match result of the first target content is not a successful match, i.e., no match, then S310 is executed.
[0081] S307: If the matching result of the previous first content is a successful match, obtain the first position of the second content that successfully matches the previous first content in the second content set, wherein the position of the second content that successfully matches the current first target content in the second content set is used as the second position, and a successful match includes the two being the same or equivalent matches.
[0082] In one implementation, each piece of second content in the second content set corresponds to a second sequence number. When the matching result between the first target content and the second target content is a successful match, the second sequence number corresponding to the second target content is recorded, and the second sequence number corresponding to the first content is also recorded as the matching position of the first content. Thus, within this record, the matching content of the previous first content of the current first target content can be determined, that is, the first position of the second content matching the previous first target content within the second content set. Then, the second sequence number of the second content matching the current first target content within the second content set is used as the second position.
[0083] Additionally, it should be noted that the preceding target content refers to the first content corresponding to the first sequence number preceding the first target content in the first content set, and is not necessarily the first content that was successfully matched previously.
[0084] As one implementation method, based on the preset matching order, the preceding first content of the current first target content in the first content set is determined. Specifically, the preset matching order is the order of the first sequence number from smallest to largest. Then, it is determined whether the matching result of the preceding first content is a successful match. If it is a non-match, it is determined that the first position cannot be found, and S310 is executed directly. If it is a successful match, the first position of the second content that successfully matches the preceding first target content in the second content set is determined.
[0085] S308: Determine whether the first position and the second position are adjacent.
[0086] Considering that during the matching process between the first content set and the second content set, if consecutive second content in the second content set is successfully matched by consecutive first content in the first content set, the matching degree should be increased so that the matching degree of consecutive successful matches is greater than the matching degree of non-consecutive matches.
[0087] Therefore, if the first position and the second position are adjacent, it can be determined that after a certain first content in the first content set is successfully matched with a certain second content in the second content set, the first content after that first content in the first content set is also successfully matched with the second content adjacent to that second content in the second content set, thus achieving the effect of continuous successful matching.
[0088] S309: Increase the initial matching degree of the current first target content to obtain the matching degree of the first target content.
[0089] One implementation method is to add the initial matching degree of the first target content to a first value, and use the result of the sum as the matching degree of the first target content. The first value can be set according to actual usage requirements.
[0090] For example, assuming the first value is 1, the first index of the current first target content in the first content set is 3, and the second index of the second content in the second content set that matches the current first target content is 5, that is, the second position is 5. Assuming the matching result of the current first target content is that the two are the same, the initial matching degree is 1. If the matching result of the first content in the first content set with the first index 2 is also the same, and the second index of the second content in the second content set that successfully matches the first content with the first index 2 is 4, that is, the first position is 4, then the first position and the second position are adjacent. Then the initial matching degree is added to the first value, and the result is 2, that is, the final matching degree of the current first target content is 2.
[0091] It is understandable that the adjacency of the first and second positions can be based on a preset matching order, where the first position precedes the second position, meaning the second index corresponding to the first position is greater than the second index corresponding to the second position, and the two second indices are adjacent. However, if the second position precedes the first position, it can be determined that the two do not satisfy the adjacency relationship, and S310 can be executed. Alternatively, the initial matching degree of the current first target content can be subtracted from the second value, and the result can be used as the final matching degree of the current first target content. Taking the second position as 5 as an example, assuming the first position is 6, then the second position precedes the first position. Assuming the second value is 1, the final matching degree of the current first target content is 0.
[0092] Additionally, it should be noted that the final match score of the current first target content will not affect the matching result of the current first target content. For example, if the final match score of the current first target content is 0, its corresponding matching result will still be a successful match.
[0093] S310: Use the initial matching degree of the first target content as the matching degree of the first target content.
[0094] In one implementation, when the matching result of the first target content is an unsuccessful match (i.e., no match), the initial matching degree of the first target content is directly used as the matching degree of the first target content. For example, if the matching result of the first target content is a no match, the matching degree corresponding to the no match is 0, and the initial matching degree of the first target content is 0, then S310 is executed, and the 0 is used as the matching degree of the first target content. In another implementation, if the matching result of the first target content is a successful match, but the matching result of the previous first content is an unsuccessful match, or the matching result of the previous first content is a successful match, but the first position and the second position are adjacent, then S310 is executed, and the matching degree corresponding to the matching result of the first target content, i.e., the initial matching degree, is used as the matching degree of the first target content.
[0095] S311: Based on a preset matching order, select the next first content from the first content set as the new first target content.
[0096] S312: Determine whether the first content set has been traversed.
[0097] If the traversal is not complete, return to execute S304; if the traversal is complete, execute S313.
[0098] S313: Determine the total matching degree of the first content set based on the matching degree of each of the first contents.
[0099] As one implementation method, the matching degree of each first content in the first content set can be accumulated, and the accumulated matching degree can be used as the total matching degree.
[0100] As another implementation, to avoid differences and inaccuracies in matching scores due to varying speech lengths, and considering continuous matching, a first quantity of the first content within the first content set is obtained. Assuming this first quantity is N, and the aforementioned first value is a, the first maximum matching score of the first content set is N*b + (N-1)*a, where b is the matching score when the matching results are the same. The matching scores of each first content are summed to obtain the total matching score. The total matching score is obtained based on the total matching score and the first maximum matching score. Specifically, the ratio of the total matching score to the first maximum matching score can be used as the total matching score. For example, the total matching score is C / (N*b + (N-1)*a). Assuming both a and b are 1, the first maximum matching score is 2N-1, and the total matching score is C / (2N-1).
[0101] As another implementation, to avoid differences and inaccuracies in matching scores due to varying lengths of the interface text sequences, and considering continuous matching, a second quantity of the second content within the second content set is obtained. Assuming this second quantity is M, and the aforementioned second value is a, the second maximum matching score of the second content set is M*b + (M-1)*a, where b is the matching score when the matching results are the same. The matching scores of each second content are summed to obtain the total matching score. The total matching score is obtained based on the total matching score and the second maximum matching score. Specifically, the ratio of the total matching score to the second maximum matching score can be used as the total matching score. For example, the total matching score is C / (M*b + (M-1)*a). Assuming both a and b are 1, the second maximum matching score is 2M-1, and the total matching score is C / (2M-1).
[0102] As another implementation method, it is necessary to consider the influence of the length of the speech text sequence and the interface text sequence on the total matching degree. The total matching degree can be obtained by weighted summation of the aforementioned C / (N*b+(N-1)*a) and C / (M*b+(M-1)*a). Specifically, the total matching degree is W1*(C / (N*b+(N-1)*a))+W2*(C / (M*b+(M-1)*a)), where W1 is the first weight and W2 is the second weight. Assuming that a and b are both 1 and W1 and W2 are both 0.5, the total matching degree is C / (2(2N-1))+C / (2(2M-1)).
[0103] Therefore, based on the foregoing embodiments, this application embodiment considers that continuous matching can increase the matching degree of the first content set, making the matching results more accurate.
[0104] Please see Figure 4 , Figure 4 An embodiment of this application provides a content matching method applied to the aforementioned electronic device. The method includes steps S401 to S410.
[0105] S401: Get the first content set and the second content set.
[0106] S402: Select a first content from the first content set as the first target content.
[0107] S403: Select a second content from the second content set as the second target content.
[0108] S404: Determine the matching result of the first target content and the second target content based on their types.
[0109] As described above, the types of the first target content and the second target content may include the pinyin type and the specified type, where the specified type may be non-Chinese characters. In some embodiments, the non-Chinese characters may be the aforementioned English and punctuation marks, etc.
[0110] Specifically, several different embodiments will be described based on the different types of the first target content and the second target content.
[0111] Considering the case of polyphonic characters, such as Figure 5 As shown, in one implementation of S404, it may include S501 to S506.
[0112] S501: When both the first target content and the second target content are of the pinyin type, if at least one of the first target Chinese character corresponding to the first target content and the second target Chinese character corresponding to the second target content has multiple pinyins, take all the pinyins corresponding to the first target Chinese character as the first pinyin and all the pinyins corresponding to the second target Chinese character as the second pinyin.
[0113] In some embodiments, determine whether the types of the first target content and the second target content are both of the pinyin type. If so, determine whether the first target Chinese character corresponding to the first target content is a polyphonic character of the aforementioned target type, and determine whether the second target Chinese character corresponding to the second target content is a polyphonic character of the aforementioned target type. That is, determine whether the Chinese character corresponding to the first target content is a polyphonic character and includes multiple different pinyins, and determine whether the Chinese character corresponding to the second target content is a polyphonic character and includes multiple different pinyins. If at least one of the first target Chinese character corresponding to the first target content and the second target Chinese character corresponding to the second target content is a polyphonic character of the target type, that is, corresponds to multiple pinyins, take all the pinyins corresponding to the polyphonic character as the first pinyin or the second pinyin of the polyphonic character. For example, "重" in the aforementioned speech text sequence is a polyphonic character of the target type, and its corresponding first pinyins are chong and zhong.
[0114] S502: Determine whether there are the same pinyins between all the first pinyins and all the second pinyins.
[0115] If there are the same pinyins between all the first pinyins and all the second pinyins, it is very likely that the Chinese characters corresponding to the first pinyin and the second pinyin are the same Chinese character, and the user's pronunciation may be different from the pronunciation of this Chinese character recognized by the electronic device.
[0116] S503: Determine that the matching result of the first target content and the second target content is a polyphonic match.
[0117] Therefore, when there are identical pinyin among all the first pinyin and all the second pinyin, S503 is executed, that is, the matching result of the first target content and the second target content is determined to be a polyphonic match. In some embodiments, even if the first target content does not match all the second content, the second content that has not undergone a matching operation is no longer matched, and the matching result of the first target content is a polyphonic match.
[0118] In some embodiments, if there are no identical pinyin among all the first pinyin and all the second pinyin, the matching result of the first target content and the second target content can be determined to be a mismatch. If there are still second contents in the second content set that have not yet been matched with the first target content, the first target content continues to match with new second target contents until the first target content is matched with each of the second contents in the second content set or a second content that successfully matches the first target content is found. The successful match includes the two being identical or equivalent.
[0119] In other embodiments, if there are no identical pinyin among all the first pinyin and all the second pinyin, it can also be determined whether the first content and the second content are mixed, thereby avoiding the user's accent or inaccurate pronunciation from causing the user to mispronounce the sound, i.e., executing S504.
[0120] S504: Determine whether the first target content and the second target content are in a mixing relationship.
[0121] In some embodiments, a mixed-sound relationship refers to two pinyin syllables that are easily confused. Specifically, this can be achieved by statistically analyzing the probability of confusion between different pinyin syllables, and grouping pinyin syllables with a probability greater than a preset threshold into a pinyin group. If the first pinyin syllable and the second pinyin syllable happen to belong to this pinyin group, then they are determined to be in a mixed-sound relationship. As shown in Table 1 above, the specific implementation method for determining the mixed-sound relationship can be referred to the foregoing and will not be repeated here.
[0122] S505: Determine that the matching result between the first target content and the second target content is a mixed audio matching.
[0123] S506: Determine that the matching result between the first target content and the second target content is not a match.
[0124] Therefore, when at least one of the first target Chinese characters corresponding to the first target content and the second target Chinese characters corresponding to the second target content has multiple pinyin, after performing a matching operation on the first target content and the second target content, the matching result of the first target content can be one of three types: multi-pinyin matching, mixed-pinyin matching, or no matching. Then, S405 and subsequent operations are executed.
[0125] If both the first and second target contents are monosyllabic, such as Figure 6 As shown, in one embodiment of S404, S601 to S606 may be included.
[0126] S601: If both the first target content and the second target content are of the Pinyin type, and if the first target Chinese character corresponding to the first target content and the second target Chinese character corresponding to the second target content are both monosyllabic characters, determine whether the first target content and the second target content are consistent.
[0127] According to the pronunciation of Chinese characters, Chinese characters can be divided into polyphonic characters and monophonic characters. Monophonic characters refer to Chinese characters that have only one pronunciation or include multiple pronunciations but correspond to only one pinyin. That is to say, although there are multiple pronunciations, different pronunciations correspond to the same pinyin but only the tones are different. Therefore, in the embodiments of this application, monophonic characters refer to the meaning of a single pinyin.
[0128] S602: Determine that the matching results of the first target content and the second target content are the same.
[0129] If the Chinese character corresponding to the first target content is a monosyllabic character, and the Chinese character corresponding to the second target content is also a monosyllabic character, then it is determined whether the first target content and the second target content are consistent. If they are consistent, it means that the pronunciations of the two Chinese characters are the same, and thus it can be determined that the first pinyin and the second pinyin are consistent, that is, the matching result of the first target content and the second target content is the same.
[0130] S603: Determine whether the first content and the second content are in a mixed audio relationship.
[0131] If the Chinese character corresponding to the first target content is a monosyllabic character, and the Chinese character corresponding to the second target content is also a monosyllabic character, and the first target content and the second target content are inconsistent, it can be indicated that the two are not the same pinyin, that is, the pronunciations are different. However, considering that the user's pronunciation may not be standard enough, resulting in pronunciation errors, it can be determined whether the two are mixed. The specific judgment method can be referred to the aforementioned S504 and the previous implementation methods, which will not be repeated here.
[0132] S604: Determine that the matching result of the first target content and the second target content is a mixed audio matching.
[0133] S605: The matching result between the first target content and the second target content is determined to be a mismatch.
[0134] The situation where neither the first target content nor the second target content is entirely composed of pinyin can be divided into two cases: the first case is where one of the first and second target contents is a specified type and the other is pinyin; the second case is where both the first and second target contents are specified types. Specifically, the first case could be where the first target content is a specified type and the second target content is pinyin, or vice versa. In this embodiment, if one of the first and second target contents is a specified type and the other is pinyin, it can be directly determined that there is a mismatch between the first and second target contents, i.e., the matching result of the first target content is a mismatch. Then, if the first target content does not match all the second content in the second content set, the second content that did not match the first target content is selected and matched with the first target content again, until all the second content has been matched with the first target content or a second content that successfully matches the first target content is found.
[0135] For cases where both the first and second target content are of a specified type, such as... Figure 7 As shown, in one embodiment of S404, S701 to S703 may be included.
[0136] S701: If both the first target content and the second target content are of a specified type, determine that the first target content and the second target content are consistent.
[0137] S702: Determine that the matching result between the first target content and the second target content is not a match.
[0138] S703: Determine that the matching results of the first target content and the second target content are the same.
[0139] In one implementation, the specified type is a non-Chinese character type of text. Specifically, it can include at least one of English and punctuation marks. That is, the specified type can be English, punctuation marks, or both. If both the first target content and the second target content are English, then if the two English letters are identical, it can be determined that the first target content and the second target content are identical, that is, the English letters and their arrangement are identical. In one implementation, the English letters are not case-sensitive; that is, uppercase A and lowercase a are identical. Therefore, when determining the speech text sequence and the interface text sequence, the English letters in the sequence can be uniformly converted to lowercase letters. If both the first target content and the second target content are punctuation marks, then regardless of the type of punctuation mark, as long as it is a punctuation mark, the first target content and the second target content are determined to be identical. If one of the first target content and the second target content is English and the other is punctuation marks, then they can be determined to be inconsistent.
[0140] Therefore, the matching result of the first target content and the second target content can be determined through the above method. If the matching result is the same or equivalent, the matching result of the first target content and the second target content is determined to be a successful match, that is, the matching result of the first target content is a successful match.
[0141] S405: Determine whether the matching result of the first target content is a successful match.
[0142] If the matching result of the first target content and the second target content is a successful match, then the continuous matching needs to be considered in order to adjust the matching degree of the first target content. That is, if the matching result of the first target content is a successful match, then S406 is executed; if the matching result of the first target content is a failed match, that is, no match, then S407 is executed.
[0143] S406: Based on the continuity of the matching results, set the matching degree of the first target content.
[0144] As one implementation method, the implementation of S406 can refer to the aforementioned S307 to S310, and will not be repeated here.
[0145] S407: Determine whether the first target content has been matched with all second content.
[0146] If the first target content and the second target content fail to match, and if the first target content has not yet been traversed with all the second content in the second content set, traversal can continue, thus requiring execution of S407. If the first target content has not been matched with all the second content, return to S403, where a new second content is selected from the second content set as the new second target content. Specifically, this can be based on a preset matching order, selecting the second content after the current second target content from the second content set as the new second target content, and then executing S404 and subsequent operations. If the first target content has been traversed with all the second content in the second content set, and the matching result of the first target content is still unmatched, then execution of S408 can be performed, that is, determining that the matching result of the first target content is unmatched, and setting the matching degree of the first target content.
[0147] S408: Determine that the matching result of the first target content is not a match, and set the matching degree of the first target content.
[0148] S409: Iterate through all first content items and obtain the matching degree of each first content item.
[0149] S410: Determine the total matching degree of the first content set based on the matching degree of each of the first contents.
[0150] Therefore, the matching degree of each first content in the first content set is determined based on the type of the first content and the second content, as well as the continuity of the matching results, so as to obtain the total matching degree of the first content set, making the acquisition of the total matching of the first content set more accurate.
[0151] Please see Figure 8 , Figure 8 An embodiment of this application provides a content matching method applied to the aforementioned electronic device. The method includes steps S801 to S819.
[0152] S801: Get the first content set and the second content set.
[0153] S802: Select a first content from the first content set as the first target content.
[0154] S803: Select a second content from the second content set as the second target content.
[0155] S804: Determine whether the content of the first target and the content of the second target are both Pinyin.
[0156] S805: Determine whether there are multiple sounds in the content of the first target and the content of the second target.
[0157] If both the first and second target contents are pinyin, determine whether they are monosyllabic characters. If they are not monosyllabic characters, meaning both the first and second target contents have multiple pronunciations, then execute S806. If they are monosyllabic characters, meaning neither the first nor the second target contents have multiple pronunciations, then execute S812.
[0158] S806: Determine whether there are any identical pinyin among all the first pinyin and all the second pinyin.
[0159] S807: Determine that the matching result between the first target content and the second target content is a polyphonic matching.
[0160] S808: Determine whether the first content and the second content are in a mixed audio relationship.
[0161] S809: Determine that the matching result of the first target content and the second target content is a mixed audio matching.
[0162] S810: Determine that the matching result between the first target content and the second target content is not a match.
[0163] S811: Determine whether the content of the first target and the content of the second target are both of the specified type.
[0164] If neither the first target content nor the second target content is Pinyin, then it is determined whether both the first target content and the second target content are of the specified type. If so, S812 is executed. If neither the first target content nor the second target content is of the specified type, then neither the first target content nor the second target content is of the specified type nor is both Pinyin. Therefore, one of the first target content and the second target content is Pinyin and the other is of the specified type. In this case, S810 is executed, and the matching result between the first target content and the second target content is determined to be a mismatch.
[0165] S812: Determine whether the content of the first target is consistent with the content of the second target.
[0166] S813: Determine that the matching results of the first target content and the second target content are the same.
[0167] S814: Determine whether the matching result of the first target content is a successful match.
[0168] S815: Based on the continuity of the matching results, set the matching degree of the first target content.
[0169] S816: Determine whether the first target content has been matched with all the second content.
[0170] S817: Determine that the matching result of the first target content is not a match, and set the matching degree of the first target content.
[0171] S818: Iterate through all first content items and obtain the matching degree of each first content item.
[0172] S819: Determine the total matching degree of the first content set based on the matching degree of each of the first contents.
[0173] It should be noted that S801 to S819 can be referred to the aforementioned embodiments, and will not be repeated here.
[0174] Please see Figure 9 The diagram shows a structural block diagram of a content matching device 900 provided in an embodiment of this application. The device may include: a first acquisition unit 901, a second acquisition unit 902, a determination unit 903, a statistics unit 904, and a calculation unit 905.
[0175] The first acquisition unit 901 is used to acquire a first content set, which is acquired in advance based on the user's input voice. The first content set includes at least one first content, and the at least one first content includes the first pinyin of the first Chinese character in the voice.
[0176] The second acquisition unit 902 is used to acquire a second content set, the second content set including at least one second content of target description information, the at least one second content including the second pinyin of the second Chinese character corresponding to the target description information, and the target description information corresponding to the target operation.
[0177] The determining unit 903 is configured to perform a matching operation on each of the first content and the second content set, and determine the matching result for each of the first content, wherein the matching result includes the two being the same, equivalent matching, and non-matching, and the equivalent matching includes at least one of mixed sound matching and multi-sound matching.
[0178] Furthermore, equivalent matching includes mixed sound matching, which is used to characterize that the first content and the second content belong to a mixed sound relationship, and the mixed sound relationship includes at least one of front and back nasal sounds, nasal lateral sounds, and retroflex consonant sounds.
[0179] Furthermore, the equivalent matching includes polyphonic matching, which is used to characterize that at least one of the first Chinese character corresponding to the first content and the second Chinese character corresponding to the second content corresponds to multiple pinyin, and all the pinyin corresponding to the first Chinese character is the first pinyin, all the pinyin corresponding to the second Chinese character is the second pinyin, and there are the same pinyin in the first pinyin and the second pinyin.
[0180] Furthermore, the determining unit 903 is also configured to select a first content from the first content set as a first target content; match the first target content with the second content set to obtain a matching result of the first target content; determine the matching degree of the first target content based on the matching result of the first target content; and select the next first content from the first content set as a new first target content based on a preset matching order, so as to traverse each first content in the first content set and determine the matching degree of each first content.
[0181] Furthermore, the determining unit 903 is also configured to select a second content from the second content set as the second target content; when both the first target content and the second target content are of the pinyin type, if at least one of the first target Chinese characters corresponding to the first target content and the second target Chinese characters corresponding to the second target content has multiple pinyin, all the pinyin corresponding to the first target Chinese characters are taken as the first pinyin and all the pinyin corresponding to the second target Chinese characters are taken as the second pinyin; if there is a common pinyin among all the first pinyin and all the second pinyin, the matching result of the first target content and the second target content is determined to be a polyphonic matching; if there is no common pinyin among all the first pinyin and all the second pinyin, the matching result of the first target content and the second target content is determined to be a mismatch, and each second content is traversed and a matching operation is performed with the first target content until the first target content is matched with each second content in the second content set or a second content that successfully matches the first target content is found, wherein the successful matching includes matching the same or equivalent content.
[0182] Furthermore, the determining unit 903 is also used to determine whether the first content and the second content belong to a mixed sound relationship if there is no common pinyin among all the first pinyin and all the second pinyin. The mixed sound relationship includes at least one of front and back nasal sounds, nasal and lateral sounds, and retroflex and alveolar sounds. If they belong to a mixed sound relationship, the matching result of the first target content and the second target content is determined to be a mixed sound match. If they do not belong to a mixed sound relationship, the matching result of the first target content and the second target content is determined to be a mismatch.
[0183] Furthermore, the determining unit 903 is also used to select a second content as the second target content from the first content set; when both the first target content and the second target content are of the pinyin type, if the first target Chinese character corresponding to the first target content and the second target Chinese character corresponding to the second target content are both monosyllabic characters, it is determined whether the first target content and the second target content are consistent; if they are consistent, it is determined that the matching result of the first target content and the second target content is the same; if they are inconsistent, it is determined that the matching result of the first target content and the second target content is not a match, and each second content is traversed and matched with the first target content until the first target content is matched with each second content in the second content set or a second content that successfully matches the first target content is found, wherein the successful match includes the same or equivalent match.
[0184] Furthermore, the determining unit 903 is also used to determine whether the first content and the second content belong to a mixed sound relationship if they are inconsistent. The mixed sound relationship includes at least one of front and back nasal sounds, nasal lateral sounds, and retroflex consonant sounds. If they belong to a mixed sound relationship, the matching result of the first target content and the second target content is determined to be a mixed sound match. If they do not belong to a mixed sound relationship, the matching result of the first target content and the second target content is determined to be a mismatch.
[0185] Furthermore, the determining unit 903 is also configured to select a second content as the second target content from the first content set; when both the first target content and the second target content are of a specified type, if the first target content and the second target content are consistent, then the matching result of the first target content and the second target content is determined to be the same; if the first target content and the second target content are inconsistent, then the matching result of the first target content and the second target content is determined to be a mismatch, and each second content is traversed and matched with the first target content until the first target content is matched with each second content in the second content set or a second content that successfully matches the first target content is found, wherein the successful match includes the two being the same or equivalent matches.
[0186] Furthermore, the determining unit 903 is also used to select a second content as the second target content from the first content set; if one of the first target content and the second target content is a specified type and the other is a pinyin type, then the matching result of the first target content and the second target content is determined to be a mismatch, and each second content is traversed and matched with the first target content until the first target content is matched with each second content in the second content set or a second content that successfully matches the first target content is found, wherein the successful match includes the two being the same or equivalent matches.
[0187] Furthermore, the determining unit 903 is also configured to determine the initial matching degree of the first target content based on the matching result of the first target content; if the matching result of the first target content is a successful match, determine the preceding first content of the current first target content in the first content set based on the preset matching order; if the matching result of the preceding first content is a successful match, obtain the first position of the second content that successfully matches the preceding first content in the second content set, wherein the position of the second content that successfully matches the current first target content in the second content set is used as the second position, and successful matching includes the two being the same or equivalent matches; if the first position and the second position are adjacent, increase the initial matching degree of the current first target content to obtain the matching degree of the first target content; if the matching result of the first target content is a non-match, use the initial matching degree of the first target content as the matching degree of the first target content.
[0188] The statistics unit 904 is used to determine the matching degree of the first content based on the matching results of the first content, and the matching degree is different for different matching results.
[0189] The calculation unit 905 is used to determine the total matching degree of the first content set based on the matching degree of each of the first contents, and the total matching degree is used as the matching result between the first content set and the second content set.
[0190] Furthermore, the calculation unit 905 is also used to obtain a first number of first contents within the first content set; to sum the matching degree of each first content to obtain a total matching degree; and to obtain the total matching degree based on the total matching degree and the first number.
[0191] Furthermore, the calculation unit 905 is also used to obtain a second quantity of the second content within the second content set; obtain a first matching degree based on the total matching degree and the first quantity; obtain a second matching degree based on the total matching degree and the second quantity; and obtain the total matching degree based on the first matching degree and the second matching degree.
[0192] Among them, the matching degree corresponding to the same, equivalent matching and non-matching decreases in that order.
[0193] Furthermore, the content matching device 900 also includes a processing unit for performing the target operation if the total matching degree meets a preset condition.
[0194] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0195] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0196] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0197] Please refer to Figure 10 This document illustrates a structural block diagram of an electronic device according to an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications can be stored in the memory 120 and configured to be executed by one or more processors 110, and the one or more applications are configured to perform the methods described in the foregoing method embodiments.
[0198] Processor 110 may include one or more processing cores. Processor 110 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data of the electronic device 100 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.
[0199] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the terminal 100 during use (such as phonebook data, audio and video data, chat log data, etc.).
[0200] Please refer to Figure 11 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 1100 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0201] The computer-readable storage medium 1100 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 1100 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1100 has storage space for program code 1110 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 1110 may, for example, be compressed in a suitable form.
[0202] In summary, the content matching method, apparatus, electronic device, computer-readable medium, and computer program product provided in this application perform a matching operation on each first content set and the second content set to determine the matching result of each first content set. Since the matching result includes identical, equivalent, and non-matching, and the equivalent matching includes at least one of mixed-sound matching and multi-sound matching, compared to simply determining identical or different matching results, it can take into account the effects of mixed and multi-sound matching in pinyin, so that the matching result also includes at least one of mixed-sound matching and multi-sound matching. Therefore, when determining the matching degree of each first content set based on the matching result and then determining the total matching degree of the first content set, the determination of the total matching degree can be more accurate, thereby improving the accuracy of matching the relationship between speech recognition content and target description information.
[0203] The beneficial effects of the embodiments of this application are as follows:
[0204] 1. Pinyin-based matching reduces problems caused by inaccurate user pronunciation or low recognition rate of homophones in speech recognition. For example, if a user wants to say "Chengdu" but the speech recognition says "Chengdu", conventional Chinese character matching or semantic matching will not be able to match it, but pinyin-based matching can match it well.
[0205] 2. Added fuzzy matching processing for polyphonic and mixed-sound characters, which reduces the problem of inaccurate recognition caused by users' inaccurate pronunciation of nasal sounds or low speech recognition rate. For example, the problem of users wanting to say "support" but the speech recognition is "knowledge" can be solved by adding mixed-sound character processing.
[0206] 3. The matching algorithm is improved by adding a continuity dimension to the matching degree calculation. The matching degree is calculated from multiple dimensions, which increases the accuracy of the matching algorithm. This is because the continuity of the text sequence matching also represents the degree of matching. For example, the matching degree of "Chengdu" and "Chengdu City" should be higher than that of "Chengdu" and "Chengdu".
[0207] 4. The matching algorithm calculates the matching degree in both directions based on the speech-text sequence and the interface text sequence, which can effectively solve the problem of balancing the difference in matching degree calculated based on two different sequences. For example, the matching degree of "Chengdu" and "Chengdu City" is different from the matching degree of "Chengdu" and "Chengdu High-tech Zone". The matching degree of "Chengdu" and "Chengdu City" should be greater.
[0208] 5. Applicable to most scenarios involving fuzzy matching of Chinese characters and other texts after speech recognition, requiring no persistent storage, consuming less memory, and involving no database operations.
[0209] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A content matching method characterized by, The method comprises: obtaining a first content set, the first content set being obtained in advance based on a voice input by a user, the first content set comprising at least one first content, the at least one first content comprising a first pinyin of a first Chinese character in the voice; obtaining a second content set, the second content set comprising at least one second content of target description information, the at least one second content comprising a second pinyin of a second Chinese character corresponding to the target description information, the target description information corresponding to a target operation; selecting one first content from the first content set as a first target content; matching the first target content with the second content set to obtain a matching result of the first target content; determining an initial matching degree of the first target content based on the matching result of the first target content; if the matching result of the first target content is successful matching, determining a previous first content of the first target content in the first content set based on a preset matching order; if the matching result of the previous first content is successful matching, obtaining a first position of the second content successfully matched with the previous first content in the second content set, wherein a position of the second content successfully matched with the first target content in the second content set is a second position, successful matching includes identical or equivalent matching, and the equivalent matching includes at least one of mixed tone matching and multi-tone matching; if the first position and the second position are adjacent, increasing the initial matching degree of the first target content to obtain a matching degree of the first target content, wherein the first position and the second position being adjacent includes the first position being a position before the second position; if the first position and the second position are not adjacent, subtracting a second value from the initial matching degree of the first target content to obtain a result as the matching degree of the first target content, wherein the first position and the second position not being adjacent includes the second position being a position before the first position; if the matching result of the first target content is not matching, taking the initial matching degree of the first target content as the matching degree of the first target content; based on the preset matching order, selecting a next first content from the first content set as a new first target content to traverse each first content in the first content set and determine a matching degree of each first content; determining a total matching degree of the first content set based on the matching degree of each first content, the total matching degree being a matching result between the first content set and the second content set.
2. The method of claim 1, wherein, After determining the total matching degree of the first content set based on the matching degree of each first content, the method further comprises: if the total matching degree satisfies a preset condition, performing the target operation.
3. The method of claim 1, wherein, The equivalent matching includes mixed tone matching, and the mixed tone matching is used to represent a mixed tone relationship between the first content and the second content, the mixed tone relationship including at least one of front and back nasal tone, nasal tone side tone, and flat and raised tongue tone.
4. The method of claim 1, wherein, The equivalent matching includes multi-pinyin matching, the multi-pinyin matching is used for representing that at least one of the first Chinese character corresponding to the first content and the second Chinese character corresponding to the second content corresponds to multiple pinyins, all pinyins corresponding to the first Chinese character are first pinyins, all pinyins corresponding to the second Chinese character are second pinyins, and there is the same pinyin in the first pinyins and the second pinyins.
5. The method of claim 1, wherein, The matching of the first target content with the second content set to obtain the matching result of the first target content comprises: selecting one second content from the second content set as a second target content; in the case that the first target content and the second target content are both pinyin types, if at least one of the first target Chinese character corresponding to the first target content and the second target Chinese character corresponding to the second target content corresponds to multiple pinyins, all pinyins corresponding to the first target Chinese character are first pinyins and all pinyins corresponding to the second target Chinese character are second pinyins; if there is the same pinyin between all the first pinyins and all the second pinyins, it is determined that the matching result of the first target content and the second target content is multi-pinyin matching; if there is no same pinyin between all the first pinyins and all the second pinyins, it is determined that the matching result of the first target content and the second target content is not matching, and the matching operation is performed between each second content and the first target content until the matching of the first target content with each second content in the second content set is completed or the second content successfully matched with the first target content is found, and the successful matching includes same or equivalent matching.
6. The method of claim 5, wherein, if there is no same pinyin between all the first pinyins and all the second pinyins, it is determined that the matching result of the first target content and the second target content is not matching, which comprises: if there is no same pinyin between all the first pinyins and all the second pinyins, it is determined whether the first content and the second content belong to a mixed pinyin relationship, the mixed pinyin relationship includes at least one of front and back nasal sounds, nasal sound side sound and flat and raised tongue sound; if it belongs to the mixed pinyin relationship, it is determined that the matching result of the first target content and the second target content is mixed pinyin matching; if it does not belong to the mixed pinyin relationship, it is determined that the matching result of the first target content and the second target content is not matching.
7. The method of claim 1, wherein, The matching of the first target content with the second content set to obtain the matching result of the first target content comprises: selecting one second content from the second content set as a second target content; in the case that the first target content and the second target content are both pinyin types, if the first target Chinese character corresponding to the first target content and the second target Chinese character corresponding to the second target content are both single-pinyin characters, it is determined whether the first target content and the second target content are consistent; if they are consistent, it is determined that the matching result of the first target content and the second target content is same; If the first target content and the second target content are inconsistent, it is determined that the matching result of the first target content and the second target content is not matched, and a matching operation is performed on each second content and the first target content until the first target content is matched with each second content in the second content set or a second content successfully matched with the first target content is found, and the successful matching includes identical or equivalent matching.
8. The method of claim 7, wherein, If the first target content and the second target content are inconsistent, it is determined that the matching result of the first target content and the second target content is not matched, and a matching operation is performed on each second content and the first target content until the first target content is matched with each second content in the second content set or a second content successfully matched with the first target content is found, and the successful matching includes identical or equivalent matching. If the first target content and the second target content are inconsistent, it is determined that the matching result of the first target content and the second target content is not matched, and a matching operation is performed on each second content and the first target content until the first target content is matched with each second content in the second content set or a second content successfully matched with the first target content is found, and the successful matching includes identical or equivalent matching. If the first target content and the second target content are inconsistent, it is determined that the matching result of the first target content and the second target content is not matched, and a matching operation is performed on each second content and the first target content until the first target content is matched with each second content in the second content set or a second content successfully matched with the first target content is found, and the successful matching includes identical or equivalent matching. The at least one first content further includes a first character, and the at least one second content further includes a second character, the first character and the second content are of a specified type, the specified type is a non-Hanzi type character, and the matching of the first target content with the second content set to obtain the matching result of the first target content includes:
9. The method of claim 1, wherein, selecting one second content from the first content set as a second target content; If the first target content and the second target content are inconsistent, it is determined that the matching result of the first target content and the second target content is not matched, and a matching operation is performed on each second content and the first target content until the first target content is matched with each second content in the second content set or a second content successfully matched with the first target content is found, and the successful matching includes identical or equivalent matching. The at least one first content further includes a first character, and the at least one second content further includes a second character, the first character and the second content are of a specified type, the specified type is a non-Hanzi type character, and the matching of the first target content with the second content set to obtain the matching result of the first target content includes: selecting one second content from the first content set as a second target content; 10. The method of claim 1, wherein, If one of the first target content and the second target content is of the specified type and the other is of the pinyin type, it is determined that the matching result of the first target content and the second target content is not matched, and a matching operation is performed on each second content and the first target content until the first target content is matched with each second content in the second content set or a second content successfully matched with the first target content is found, and the successful matching includes identical or equivalent matching. 11. The method according to any of claims 1 to 10, characterized in that The total matching degree of the first content set is determined based on the matching degree of each of the first content, comprising: obtaining a first quantity of the first content in the first content set; adding the matching degree of each of the first content to obtain a total matching degree; obtaining the total matching degree based on the total matching degree and the first quantity.
12. The method of claim 11, wherein, The total matching degree is obtained based on the total matching degree and the first quantity, comprising: obtaining a second quantity of the second content in the second content set; obtaining a first matching degree based on the total matching degree and the first quantity; obtaining a second matching degree based on the total matching degree and the second quantity; obtaining the total matching degree according to the first matching degree and the second matching degree.
13. The method according to any one of claims 1 to 10, characterized in that, The matching degrees corresponding to the same, equivalent matching and no matching are reduced in turn.
14. A content matching apparatus characterized by comprising: Comprising: a first obtaining unit, configured to obtain a first content set, the first content set being obtained based on a voice input by a user in advance, the first content set comprising at least one first content, the at least one first content comprising a first pinyin of a first Chinese character in the voice; a second obtaining unit, configured to obtain a second content set, the second content set comprising at least one second content of target description information, the at least one second content comprising a second pinyin of a second Chinese character corresponding to the target description information, the target description information corresponding to a target operation; a determining unit, configured to select one first content from the first content set as a first target content; matching the first target content with the second content set to obtain a matching result of the first target content; and determining an initial matching degree of the first target content based on the matching result of the first target content. If the matching result of the first target content is successful matching, a previous first content of the first target content in the first content set is determined based on a preset matching order; if the matching result of the previous first content is successful matching, a first position of the second content successfully matched with the previous first content in the second content set is obtained, wherein a position of the second content successfully matched with the first target content in the second content set is taken as a second position, successful matching includes identical or equivalent matching, and the equivalent matching includes at least one of mixed sound matching and multi-sound matching; if the first position and the second position are adjacent, an initial matching degree of the first target content is increased to obtain a matching degree of the first target content, wherein the first position and the second position being adjacent includes that the first position is a previous position of the second position; if the first position and the second position are not adjacent, the initial matching degree of the first target content is reduced by a second value to obtain a result as the matching degree of the first target content, wherein the first position and the second position not being adjacent includes that the second position is a previous position of the first position; if the matching result of the first target content is not matching, the initial matching degree of the first target content is taken as the matching degree of the first target content; a next first content in the first content set is selected as a new first target content based on the preset matching order, so as to traverse each first content in the first content set and determine a matching degree of each first content. A calculation unit is configured to determine a total matching degree of the first content set based on the matching degree of each first content, and the total matching degree is taken as a matching result between the first content set and the second content set.
15. An electronic device, comprising: The computer readable medium stores processor-executable program code, which, when executed by the processor, causes the processor to perform the method of any one of claims 1-13. The computer program / instruction, when executed by the processor, implements the method of any one of claims 1-13. The computer program / instruction, when executed by the processor, implements the method of any one of claims 1-13. 16. A computer readable medium characterized by 17. A computer program product comprising computer programs / instructions, characterized in that,
Citation Information
Patent Citations
Voice recognition matching method, terminal and computer readable storage medium
CN109036419A
Matching method and system for site or person
CN111627445A