Dialect recognition model construction method and device, storage medium and electronic device

CN116386597BActive Publication Date: 2025-08-29QINGDAO HAIER TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310189393.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-01
Publication Date
2025-08-29
Estimated Expiration
2043-03-01

AI Technical Summary

Technical Problem

[0003]目前,语音识别技术在智能设备的控制等等广泛领域发挥重要作用,比如对智能家电设备进行语音控制,极大便利了人们的日常生活,大多数基于语音识别技术的语音识别模型可以较为准确地识别标准的普通话,某一些方言识别模型在构建过程中采取标准的方言语料,比如,闽南话的闽南方言识别模型在构建过程中,采用标准的闽南方言进行训练,但是,现实情况下,闽南地域辽阔,由于方言传播过程中地域阻隔等等因素,导致闽南地域除了标准的闽南方言,会派生出种类繁多的子方言,并且各个子方言之间差异或大或小,使用闽南方言识别模型对子方言进行识别常常出现识别错误的情况,造成不良的用户体验

Benefits of technology

[0042]In an embodiment of the present application, the regional center of the target dialect area is detected, wherein the dialect used in the target dialect area is the target dialect to be identified; with the regional center as the center position, the target dialect area is divided from the inside to the outside into a plurality of dialect collection intervals, wherein the plurality of dialect collection intervals are arranged radially from the inside to the outside with the regional center as the center, and the closer the dialect collection interval is to the regional center, the higher the similarity between the language used and the target dialect; the dialect speech expressing the target semantics in each dialect collection interval of the plurality of dialect collection intervals is collected to obtain a plurality of sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships; the plurality of sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships are used to train an initial dialect recognition model to obtain the target dialect. A target dialect recognition model is provided, wherein the target dialect recognition model is used to recognize the target dialect, that is, before constructing the dialect recognition model corresponding to the target dialect, the regional center of the target dialect area where the target dialect is used is detected, and with the regional center as the center position, the target dialect area is divided from the inside to the outside into multiple dialect collection intervals that are arranged radially from the inside to the outside with the regional center as the center, and then the dialect speech expressing the target semantics in each dialect collection interval of the multiple dialect collection intervals is collected to obtain multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships, and finally the multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships are used to train the initial dialect recognition model to obtain the target dialect recognition model for recognizing the target dialect. This method for constructing a dialect recognition model differs from simply training a dialect recognition model using standard mainstream dialect speech. It also fully considers the relationship between dialects and regions. Specifically, dialect speech features from closely located sampling intervals are more similar. As a result, the target dialect recognition model constructed can overcome regional dialect differences and accurately identify the target dialect within the target dialect region. This technical solution addresses the low dialect recognition accuracy of dialect recognition models in related technologies, achieving the technical effect of improving the dialect recognition accuracy of dialect recognition models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386597B_ABST
    Figure CN116386597B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for constructing a dialect recognition model, a storage medium, and an electronic device, and relates to the field of smart home technology. The method for constructing the dialect recognition model includes: detecting the regional center of a target dialect area; dividing the target dialect area from the inside to the outside into multiple dialect collection intervals with the regional center as the center position; collecting dialect speech expressing target semantics in each of the multiple dialect collection intervals to obtain multiple sets of dialect collection intervals, target semantics, and dialect speech with corresponding relationships; using the multiple sets of dialect collection intervals, target semantics, and dialect speech with corresponding relationships to train an initial dialect recognition model to obtain a target dialect recognition model. The above technical solution solves the problem of low dialect recognition accuracy of the dialect recognition model in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart home technology, and more specifically, to a method and device for constructing a dialect recognition model, a storage medium, and an electronic device. Background Art

[0002] Speech recognition technology, also known as automatic speech recognition (ASR), aims to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. This differs from speaker identification and speaker verification, which attempt to identify or verify the speaker of the speech rather than the lexical content.

[0003] At present, speech recognition technology plays an important role in a wide range of fields such as the control of smart devices. For example, voice control of smart home appliances has greatly facilitated people's daily lives. Most speech recognition models based on speech recognition technology can recognize standard Mandarin relatively accurately. Some dialect recognition models use standard dialect materials during the construction process. For example, the Minnan dialect recognition model of Minnan dialect is trained using the standard Minnan dialect during the construction process. However, in reality, Minnan is a vast area. Due to factors such as geographical barriers during the spread of dialects, in addition to the standard Minnan dialect, a wide variety of sub-dialects have been derived in Minnan. The differences between the sub-dialects are large or small. Using the Minnan dialect recognition model to recognize sub-dialects often results in recognition errors, resulting in a poor user experience.

[0004] In related technologies, there is no effective solution to the problem of low dialect recognition accuracy of dialect recognition models. Summary of the Invention

[0005] The embodiments of the present application provide a method and device for constructing a dialect recognition model, a storage medium, and an electronic device to at least solve the problem of low dialect recognition accuracy of the dialect recognition model in the related art.

[0006] According to one embodiment of the present application, a method for constructing a dialect recognition model is provided, comprising:

[0007] Detecting a regional center of a target dialect region, wherein the dialect used in the target dialect region is the target dialect to be identified;

[0008] Taking the regional center as the center, the target dialect region is divided into a plurality of dialect collection intervals from the inside to the outside, wherein the plurality of dialect collection intervals are radially arranged from the inside to the outside with the regional center as the center, and the closer the dialect collection interval is to the regional center, the higher the similarity between the language used and the target dialect;

[0009] Collecting dialect speech expressing target semantics in each of the plurality of dialect collection intervals to obtain multiple sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships;

[0010] An initial dialect recognition model is trained using multiple sets of corresponding dialect collection intervals, target semantics and dialect speech to obtain a target dialect recognition model, wherein the target dialect recognition model is used to recognize the target dialect.

[0011] Optionally, the target dialect area is divided into a plurality of dialect collection intervals from the inner to the outer sides with the area center as the center position, including:

[0012] Obtaining the postal codes of all regions within the target dialect region to obtain a postal code set, wherein the central postal code of the regional center corresponding to the target dialect region is the smallest postal code in the postal code set, and the multiple postal codes in the postal code set are arranged in ascending order;

[0013] Using the central postal code as a starting value, the postal codes included in the postal code set are divided into a plurality of equally spaced numerical intervals in ascending order;

[0014] The regions corresponding to one or more postal codes in the same numerical range are divided into a dialect collection interval to obtain a plurality of dialect collection intervals.

[0015] Optionally, the step of collecting dialect speech expressing the target semantics in each of the plurality of dialect collection intervals to obtain a plurality of sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships includes:

[0016] Sending text data expressing the target semantics to terminal devices located in each of the dialect collection intervals;

[0017] The voice data returned by the terminal device in response to the text data is received as the dialect voice, and a plurality of sets of dialect collection intervals, target semantics and dialect voices having corresponding relationships are obtained.

[0018] Optionally, the step of collecting dialect speech expressing the target semantics in each of the plurality of dialect collection intervals to obtain a plurality of sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships includes:

[0019] Identifying speech data expressing the target semantics from the speech database corresponding to each of the dialect collection intervals as the dialect speech;

[0020] Construct multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships.

[0021] Optionally, the method of using multiple sets of corresponding dialect collection intervals, target semantics, and dialect speech to train an initial dialect recognition model to obtain a target dialect recognition model includes:

[0022] Construct the initial dialect recognition model as the initial current dialect recognition model, and repeat the following steps until the target dialect recognition model is obtained:

[0023] According to the arrangement order of the dialect collection intervals in the target dialect area from inside to outside, the dialect collection intervals are sequentially acquired as current intervals;

[0024] Training the current dialect recognition model using the target semantics and dialect speech corresponding to the current interval to obtain a first dialect recognition model;

[0025] If the current interval is not the last interval in the arrangement order, determining the first dialect recognition model as the next current dialect recognition model;

[0026] In a case where the current interval is the last interval in the arrangement order, the first dialect recognition model is determined as the target dialect recognition model.

[0027] Optionally, the method of using multiple sets of corresponding dialect collection intervals, target semantics, and dialect speech to train an initial dialect recognition model to obtain a target dialect recognition model includes:

[0028] Obtaining a target number of initial dialect recognition models, wherein the target number is the number of dialect collection intervals;

[0029] Using a set of corresponding target semantics and dialect speech to train an initial dialect recognition model, respectively, to obtain the target number of second dialect recognition models;

[0030] averaging the model parameters of the target number of second dialect recognition models to obtain target model parameters;

[0031] The initial dialect recognition model having the target model parameters is determined as the target dialect recognition model.

[0032] Optionally, detecting the regional center of the target dialect region includes:

[0033] Obtaining the population density distribution of the target dialect area;

[0034] The area with the highest population density in the target dialect area is determined as the area center.

[0035] According to another embodiment of the present application, a device for constructing a dialect recognition model is provided, including:

[0036] a detection module, configured to detect a regional center of a target dialect region, wherein the dialect used in the target dialect region is the target dialect to be identified;

[0037] a division module, configured to divide the target dialect region into a plurality of dialect collection intervals from the inside out, with the regional center as the center position, wherein the plurality of dialect collection intervals are radially arranged from the inside out, with the regional center as the center, and the closer the dialect collection interval is to the regional center, the higher the similarity between the language used and the target dialect;

[0038] A collection module, configured to collect dialect speech expressing target semantics in each of the plurality of dialect collection intervals, and obtain multiple sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships;

[0039] The training module is used to train an initial dialect recognition model using multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships to obtain a target dialect recognition model, wherein the target dialect recognition model is used to recognize the target dialect.

[0040] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for constructing a dialect recognition model when running.

[0041] According to another aspect of an embodiment of the present application, an electronic device is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method for constructing the dialect recognition model through the computer program.

[0042] In an embodiment of the present application, the regional center of the target dialect area is detected, wherein the dialect used in the target dialect area is the target dialect to be identified; with the regional center as the center position, the target dialect area is divided from the inside to the outside into a plurality of dialect collection intervals, wherein the plurality of dialect collection intervals are arranged radially from the inside to the outside with the regional center as the center, and the closer the dialect collection interval is to the regional center, the higher the similarity between the language used and the target dialect; the dialect speech expressing the target semantics in each dialect collection interval of the plurality of dialect collection intervals is collected to obtain a plurality of sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships; the plurality of sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships are used to train an initial dialect recognition model to obtain the target dialect. A target dialect recognition model is provided, wherein the target dialect recognition model is used to recognize the target dialect, that is, before constructing the dialect recognition model corresponding to the target dialect, the regional center of the target dialect area where the target dialect is used is detected, and with the regional center as the center position, the target dialect area is divided from the inside to the outside into multiple dialect collection intervals that are arranged radially from the inside to the outside with the regional center as the center, and then the dialect speech expressing the target semantics in each dialect collection interval of the multiple dialect collection intervals is collected to obtain multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships, and finally the multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships are used to train the initial dialect recognition model to obtain the target dialect recognition model for recognizing the target dialect. This method for constructing a dialect recognition model differs from simply training a dialect recognition model using standard mainstream dialect speech. It also fully considers the relationship between dialects and regions. Specifically, dialect speech features from closely located sampling intervals are more similar. As a result, the target dialect recognition model constructed can overcome regional dialect differences and accurately identify the target dialect within the target dialect region. This technical solution addresses the low dialect recognition accuracy of dialect recognition models in related technologies, achieving the technical effect of improving the dialect recognition accuracy of dialect recognition models. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0045] Figure 1 1 is a schematic diagram of a hardware environment for a method for constructing a dialect recognition model according to an embodiment of the present application;

[0046] Figure 2 is a flowchart of a method for constructing a dialect recognition model according to an embodiment of the present application;

[0047] Figure 3 is a schematic diagram of a target dialect region according to an embodiment of the present application;

[0048] Figure 4 is a schematic diagram of a method for dividing dialect collection intervals according to an embodiment of the present application;

[0049] Figure 5 It is a structural block diagram of a device for constructing a dialect recognition model according to an embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0052] According to one aspect of the embodiment of the present application, a method for constructing a dialect recognition model is provided. The method for constructing a dialect recognition model is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, Figure 1 This is a hardware environment diagram of a method for constructing a dialect recognition model according to an embodiment of the present application. The above method for constructing a dialect recognition model can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal device 102 and a server 104. Figure 1As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.

[0053] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, and a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal device 102 may be, but is not limited to, a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing machine, a smart dishwasher, a smart projection device, a smart TV, a smart clothes drying rack, smart curtains, smart audio and video, a smart socket, a smart speaker, a smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purifier, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.

[0054] In this embodiment, a method for constructing a dialect recognition model is provided, which is applied to the above-mentioned device terminal. Figure 2 is a flow chart of a method for constructing a dialect recognition model according to an embodiment of the present application, such as Figure 2 As shown, the process includes the following steps:

[0055] Step S202: detecting the regional center of the target dialect region, wherein the dialect used in the target dialect region is the target dialect to be identified;

[0056] Step S204: Dividing the target dialect region into a plurality of dialect collection intervals from the inner side to the outer side with the region center as the center, wherein the plurality of dialect collection intervals are radially arranged from the inner side to the outer side with the region center as the center, and the closer the dialect collection interval is to the region center, the higher the similarity between the language used and the target dialect;

[0057] Step S206, collecting dialect speech expressing the target semantics in each of the plurality of dialect collection intervals, to obtain a plurality of sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships;

[0058] Step S208 , using multiple sets of corresponding dialect collection intervals, target semantics and dialect speech to train an initial dialect recognition model to obtain a target dialect recognition model, wherein the target dialect recognition model is used to recognize the target dialect.

[0059] Through the above steps, before constructing a dialect recognition model corresponding to the target dialect, the regional center of the target dialect region using the target dialect is detected, and with the regional center as the center position, the target dialect region is divided from the inside to the outside into multiple dialect collection intervals arranged radially from the regional center to the outside. Then, dialect speech expressing the target semantics in each of the multiple dialect collection intervals is collected to obtain multiple sets of corresponding dialect collection intervals, target semantics, and dialect speech. Finally, the multiple sets of corresponding dialect collection intervals, target semantics, and dialect speech are used to train an initial dialect recognition model to obtain a target dialect recognition model for recognizing the target dialect. The above dialect recognition model construction method is different from training a dialect recognition model using only mainstream standard dialect speech. It also fully considers the correlation between dialects and regions. That is, the dialect speech features of dialect collection intervals in similar locations are more similar. Therefore, the constructed target dialect recognition model can overcome the dialect differences between regions and accurately recognize the target dialect in the target dialect region. The above technical solution solves the problem of low dialect recognition accuracy of the dialect recognition model in related technologies, and achieves the technical effect of improving the dialect recognition accuracy of the dialect recognition model.

[0060] In the technical solution provided in step S202 above, the target dialect region is the dialect region where the target dialect is used. The size of the target dialect region is not limited and can be, but is not limited to, a provincial region, a municipal region, a county region, etc. Figure 3 is a schematic diagram of a target dialect region according to an embodiment of the present application, such as Figure 3 As shown, the target dialect region may be, but is not limited to, composed of multiple sub-regions (region 1 to region N), wherein region 1 is the regional center of the target dialect region, and the target dialect is the dialect used in the target dialect region.

[0061] Optionally, in this embodiment, due to factors such as geographical barriers during the spread of dialects, the characteristics of dialect speech between different regions are related to the geographical location, that is, the characteristics of dialect speech between regions with closer geographical locations are smaller, that is, the characteristics of dialect speech between regions with farther geographical locations are larger. Therefore, the target dialect used in the target dialect region can derive multiple sub-dialects due to the distance of the geographical location. For example, region 1 corresponds to sub-dialect 1, region 2 corresponds to sub-dialect 2, and region N corresponds to sub-dialect N. Among them, sub-dialect 1, sub-dialect 2 and sub-dialect N all belong to the target dialect. However, since the distance between region 1 and region 2 is closer than that between region 1 and region N, the dialect speech characteristics between sub-dialect 1 and sub-dialect 2 are more similar than the dialect speech characteristics between sub-dialect 1 and sub-dialect N.

[0062] In an exemplary embodiment, the regional center of the target dialect region may be detected in the following manner, but is not limited to: obtaining the population density distribution of the target dialect region; and determining the area with the highest population density in the target dialect region as the regional center.

[0063] Optionally, in this embodiment, if Figure 3 As shown, the target dialect area can be but is not limited to being composed of multiple sub-areas (region 1 to region N). To detect the regional center of the target dialect area, the population density distribution of the target dialect area can be obtained first. For example, the population density of each region from region 1 to region N is obtained to obtain the population density distribution of the target dialect area, wherein the population density can be but is not limited to being represented by population density. The area with the largest population density is obtained from the population density distribution as the regional center of the target dialect area. For example, the population density of region 1 is the largest, so region 1 is determined as the regional center of the target dialect area.

[0064] Optionally, in this embodiment, the area with the highest population density is used as the regional center of the target dialect area, and the proportion of corpus collection corresponding to the regional center can be increased. The corpus of the regional center can be used first to train the dialect recognition model to obtain the parameters of the initial stage of the recognition model. Subsequently, the corpus corresponding to the area other than the regional center of the target dialect area can be used to continue training until the parameters converge. The above training strategy can ensure that the constructed dialect recognition model can not only recognize the mainstream sub-dialects in the area around the regional center of the target dialect, but also take into account the recognition of various sub-dialects in areas farther away from the regional center.

[0065] Optionally, in this embodiment, in addition to the above-mentioned method of determining the regional center corresponding to the target dialect area according to the population density distribution of the target dialect area, the region where the regional geometric center of the target dialect area is located can also be determined as the regional center of the target dialect area.

[0066] In the technical solution provided in the above step S204, Figure 4 is a schematic diagram of a method for dividing dialect collection intervals according to an embodiment of the present application, such as Figure 4 As shown, multiple regions are distributed in the target dialect area, where region 1 is the regional center of the target dialect area. Multiple concentric circular ring areas can be constructed with region 1 as the center. Each circular ring area corresponds to a dialect collection interval. Dialect collection intervals 1 to N are arranged radially from the inside to the outside with the regional center as the center. Dialect collection interval 1 covers area A1 and area A2, dialect collection interval 2 covers area B1 and area B2, and dialect collection interval 3 covers area C1, area C2 and area C3. The dialect collection interval is a virtual interval of a range and can be generated according to actual needs. Multiple dialect collection intervals can completely cover the target dialect area or partially cover the target dialect area. The target dialect area can be divided into multiple dialect collection intervals. The division of the dialect collection intervals can be based on the area of ​​the ring or the diameter. The growth of the ring area or the growth of the diameter between two adjacent rings can be equal or non-equal.

[0067] In an exemplary embodiment, the target dialect area can be divided into multiple dialect collection intervals from the inside to the outside with the regional center as the center position in the following manner, but not limited to: obtaining the postal codes of all areas in the target dialect area to obtain a postal code set, wherein the central postal code of the regional center corresponding to the target dialect area is the smallest postal code in the postal code set, and the multiple postal codes in the postal code set are arranged in order from small to large; using the central postal code as the starting value, the multiple postal codes included in the postal code set are divided into multiple equally spaced numerical intervals in order from small to large; the areas corresponding to one or more postal codes in the same numerical interval are divided into a dialect collection interval to obtain multiple dialect collection intervals.

[0068] Optionally, in this embodiment, in addition to the above-mentioned division in a circular manner, the target dialect area can also be divided from the inside to the outside into multiple dialect collection intervals by postal code. The target dialect area includes multiple regions, and each region can be indicated by a unique postal code. The postal code set corresponding to the target dialect area includes multiple postal codes (postal code Y1 to postal code Yn). The area corresponding to the smallest postal code in the postal code set can be determined as the regional center corresponding to the target dialect area, and the multiple postal codes included in the postal code set can be divided into multiple equally spaced numerical intervals from small to large using the smallest postal code as the starting value. For example, with an interval of 10, postal code Y1 to postal code Yn can be divided into numerical interval 1 (Y1 to Y10), numerical interval 2 (Y11 to Y20),..., numerical interval n (Yn-10 to Yn), etc., and the areas corresponding to one or more postal codes in the same numerical interval can be divided into a dialect collection interval to obtain multiple dialect collection intervals.

[0069] Optionally, in this embodiment, the numerical intervals may be divided into equal intervals or non-equal intervals, and the interval values ​​may be adjusted and set according to actual needs.

[0070] In the technical solution provided in the above step S206, the target dialect area is divided into multiple dialect collection intervals, wherein the multiple dialect collection intervals are arranged radially from the inside to the outside with the regional center as the center, and the dialect speech expressing the target semantics in each interval is collected respectively, so as to obtain multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships.

[0071] In an exemplary embodiment, the dialect speech expressing the target semantics in each of the multiple dialect collection intervals can be collected, but is not limited to, in the following manner to obtain multiple sets of dialect collection intervals, target semantics, and dialect speech with corresponding relationships: sending text data expressing the target semantics to a terminal device located in each of the dialect collection intervals; receiving the speech data returned by the terminal device in response to the text data as the dialect speech, to obtain multiple sets of dialect collection intervals, target semantics, and dialect speech with corresponding relationships.

[0072] Optionally, in this embodiment, the target semantics may be, but are not limited to, selecting semantics corresponding to speech in the target dialect that has significant characteristics that are different from other dialects, or selecting semantics corresponding to speech that is frequently used in daily life.

[0073] In an exemplary embodiment, the dialect speech expressing the target semantics in each of the multiple dialect collection intervals can be collected, but is not limited to, in the following manner to obtain multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships: identifying speech data expressing the target semantics as the dialect speech from the speech library corresponding to each of the dialect collection intervals; and constructing multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships.

[0074] Optionally, in this embodiment, the voice database stores voice data used in each of the multiple dialect collection intervals, wherein the voice data is authorized by the collector, and the dialect voice can be directly recognized from the voice database.

[0075] In the technical solution provided in the above step S208, multiple sets of corresponding dialect collection intervals, target semantics and dialect speech are used to train the initial dialect recognition model. The obtained target dialect recognition model can not only identify the mainstream sub-dialects in the area around the regional center of the target dialect, but also take into account the recognition of various sub-dialects in areas farther away from the regional center.

[0076] In an exemplary embodiment, multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships can be used to train an initial dialect recognition model to obtain a target dialect recognition model, but is not limited to the following method: construct the initial dialect recognition model as the initial current dialect recognition model, and repeat the following steps until the target dialect recognition model is obtained: according to the arrangement order of the dialect collection intervals in the target dialect area from the inside to the outside, the dialect collection intervals are sequentially obtained as the current interval; the current dialect recognition model is trained using the target semantics and dialect speech with corresponding relationships corresponding to the current interval to obtain a first dialect recognition model; if the current interval is not the last interval in the arrangement order, the first dialect recognition model is determined as the next current dialect recognition model; if the current interval is the last interval in the arrangement order, the first dialect recognition model is determined as the target dialect recognition model.

[0077] Optionally, in this embodiment, multiple sets of dialect collection intervals with corresponding relationships, target semantics and dialect speech are used to train the initial dialect recognition model, and the target dialect recognition model can be obtained by adopting a step-by-step training method. For example, N sets of dialect collection intervals with corresponding relationships, target semantics and dialect speech are used to train the initial dialect recognition model. The target semantics and dialect speech with corresponding relationships can be selected in sequence according to the arrangement order of the dialect collection intervals in the target dialect area from the inside to the outside to train the current dialect recognition model, until the N sets of dialect collection intervals with corresponding relationships, target semantics and dialect speech are used up, or the model parameters of the current dialect recognition model converge.

[0078] In an exemplary embodiment, multiple sets of corresponding dialect collection intervals, target semantics and dialect speech are used to train an initial dialect recognition model to obtain a target dialect recognition model, but is not limited to the following method: obtaining a target number of initial dialect recognition models, wherein the target number is the number of dialect collection intervals; using a set of corresponding target semantics and dialect speech to train one of the initial dialect recognition models to obtain the target number of second dialect recognition models; averaging the model parameters of the target number of second dialect recognition models to obtain target model parameters; and determining the initial dialect recognition model with the target model parameters as the target dialect recognition model.

[0079] Optionally, in this embodiment, different from the above-mentioned sequential training method, a set of corresponding target semantics and dialect speech can be used to train a target number of initial dialect recognition models respectively to obtain a target number of second dialect recognition models. Finally, the model parameters of the target number of second dialect recognition models are calculated, such as by taking the mean, taking the median, etc., to obtain the target model parameters, and the initial dialect recognition model with the target model parameters is determined as the target dialect recognition model.

[0080] Optionally, in this embodiment, during the training process of the initial dialect recognition model, correction judgments can be made based on the recognition content to determine whether the original preliminary results have large deviations, error corrections can be made, and situations where the sound is correct but the text is incorrect can be handled. At the same time, semantic supplementation can be performed to reach a range that can be understood by humans or machines, ensuring that the recognition results can be used normally.

[0081] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0082] Figure 5 is a structural block diagram of a device for constructing a dialect recognition model according to an embodiment of the present application; Figure 5 As shown, including:

[0083] A detection module 502 is configured to detect a regional center of a target dialect region, wherein the dialect used in the target dialect region is the target dialect to be identified;

[0084] a division module 504 configured to divide the target dialect region into a plurality of dialect collection intervals from the inner side to the outer side, with the region center as the center, wherein the plurality of dialect collection intervals are radially arranged from the inner side to the outer side with the region center as the center, and the closer the dialect collection interval is to the region center, the higher the similarity between the language used in the dialect collection interval and the target dialect;

[0085] The collection module 506 is configured to collect dialect speech expressing the target semantics in each of the plurality of dialect collection intervals, and obtain multiple sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships;

[0086] The training module 508 is used to train an initial dialect recognition model using multiple sets of corresponding dialect collection intervals, target semantics and dialect speech to obtain a target dialect recognition model, wherein the target dialect recognition model is used to recognize the target dialect.

[0087] Through the above embodiment, before constructing a dialect recognition model corresponding to the target dialect, the regional center of the target dialect region using the target dialect is detected, and with the regional center as the center position, the target dialect region is divided from the inside to the outside into multiple dialect collection intervals arranged radially from the regional center to the outside. Then, the dialect speech expressing the target semantics in each of the multiple dialect collection intervals is collected to obtain multiple sets of dialect collection intervals with corresponding relationships, target semantics and dialect speech. Finally, the multiple sets of dialect collection intervals with corresponding relationships, target semantics and dialect speech are used to train the initial dialect recognition model to obtain a target dialect recognition model for recognizing the target dialect. The above dialect recognition model construction method is different from training the dialect recognition model using only the mainstream standard dialect speech. It also fully considers the correlation between dialects and regions. That is, the dialect speech features of dialect collection intervals with similar locations are more similar. Therefore, the constructed target dialect recognition model can overcome the dialect differences between regions and accurately recognize the target dialect in the target dialect region. The above technical solution solves the problem of low dialect recognition accuracy of the dialect recognition model in related technologies, and achieves the technical effect of improving the dialect recognition accuracy of the dialect recognition model.

[0088] In an exemplary embodiment, the partitioning module includes:

[0089] A first acquisition unit is configured to acquire the postal codes of all regions within the target dialect region to obtain a postal code set, wherein the central postal code of the regional center corresponding to the target dialect region is the smallest postal code in the postal code set, and the multiple postal codes in the postal code set are arranged in ascending order;

[0090] A first dividing unit is configured to divide the plurality of postal codes included in the postal code set into a plurality of equally spaced numerical intervals in ascending order, using the central postal code as a starting value;

[0091] The second dividing unit is used to divide the areas corresponding to one or more postal codes in the same numerical interval into a dialect collection interval to obtain multiple dialect collection intervals.

[0092] In an exemplary embodiment, the acquisition module includes:

[0093] a sending unit, configured to send text data expressing the target semantics to a terminal device located within each of the dialect collection intervals;

[0094] The receiving unit is used to receive the voice data returned by the terminal device in response to the text data as the dialect voice, and obtain multiple sets of dialect collection intervals, target semantics and dialect voices with corresponding relationships.

[0095] In an exemplary embodiment, the acquisition module includes:

[0096] a recognition unit, configured to recognize speech data expressing the target semantics from the speech library corresponding to each of the dialect collection intervals as the dialect speech;

[0097] The first construction unit is used to construct multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships.

[0098] In an exemplary embodiment, the training module includes:

[0099] The second construction unit is configured to construct the initial dialect recognition model as an initial current dialect recognition model, and repeatedly perform the following steps until the target dialect recognition model is obtained:

[0100] A second acquiring unit is configured to sequentially acquire the dialect collection intervals as current intervals according to the arrangement order of the dialect collection intervals in the target dialect region from inside to outside;

[0101] A first training unit is configured to train the current dialect recognition model using the target semantics and dialect speech corresponding to the current interval to obtain a first dialect recognition model;

[0102] a first determining unit, configured to determine the first dialect recognition model as a next current dialect recognition model if the current interval is not the last interval in the arrangement order;

[0103] The second determining unit is configured to determine the first dialect recognition model as the target dialect recognition model when the current interval is the last interval in the arrangement order.

[0104] In an exemplary embodiment, the training module includes:

[0105] a third acquiring unit, configured to acquire a target number of the initial dialect recognition models, wherein the target number is the number of the dialect collection intervals;

[0106] A second training unit is configured to train one of the initial dialect recognition models using a set of corresponding target semantics and dialect speech to obtain the target number of second dialect recognition models;

[0107] an averaging unit, configured to average the model parameters of the target number of second dialect recognition models to obtain target model parameters;

[0108] The third determining unit is configured to determine the initial dialect recognition model having the target model parameters as the target dialect recognition model.

[0109] In an exemplary embodiment, the detection module includes:

[0110] a fourth acquiring unit, configured to acquire a population density distribution of the target dialect region;

[0111] The fourth determining unit is configured to determine the area with the highest population density in the target dialect region as the regional center.

[0112] An embodiment of the present application further provides a storage medium, which includes a stored program, wherein the program executes any of the above methods when it is run.

[0113] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:

[0114] S1, detecting a regional center of a target dialect region, wherein the dialect used in the target dialect region is the target dialect to be identified;

[0115] S2, dividing the target dialect region into a plurality of dialect collection intervals from the inner side to the outer side with the regional center as the center, wherein the plurality of dialect collection intervals are radially arranged from the inner side to the outer side with the regional center as the center, and the closer the dialect collection interval is to the regional center, the higher the similarity between the language used in the dialect collection interval and the target dialect;

[0116] S3, collecting dialect speech expressing the target semantics in each of the plurality of dialect collection intervals, to obtain a plurality of sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships;

[0117] S4, using multiple sets of corresponding dialect collection intervals, target semantics and dialect speech to train an initial dialect recognition model to obtain a target dialect recognition model, wherein the target dialect recognition model is used to recognize the target dialect.

[0118] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0119] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0120] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0121] S1, detecting a regional center of a target dialect region, wherein the dialect used in the target dialect region is the target dialect to be identified;

[0122] S2, dividing the target dialect region into a plurality of dialect collection intervals from the inner side to the outer side with the regional center as the center, wherein the plurality of dialect collection intervals are radially arranged from the inner side to the outer side with the regional center as the center, and the closer the dialect collection interval is to the regional center, the higher the similarity between the language used in the dialect collection interval and the target dialect;

[0123] S3, collecting dialect speech expressing the target semantics in each of the plurality of dialect collection intervals, to obtain a plurality of sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships;

[0124] S4, using multiple sets of corresponding dialect collection intervals, target semantics and dialect speech to train an initial dialect recognition model to obtain a target dialect recognition model, wherein the target dialect recognition model is used to recognize the target dialect.

[0125] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.

[0126] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0127] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into separate integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0128] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for constructing a dialect recognition model, characterized in that: include: Detecting a regional center of a target dialect region, wherein the dialect used in the target dialect region is the target dialect to be identified; Taking the regional center as the center, the target dialect region is divided into a plurality of dialect collection intervals from the inside to the outside, wherein the plurality of dialect collection intervals are radially arranged from the inside to the outside with the regional center as the center, and the closer the dialect collection interval is to the regional center, the higher the similarity between the language used and the target dialect; Collecting dialect speech expressing target semantics in each of the plurality of dialect collection intervals to obtain multiple sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships; An initial dialect recognition model is trained using multiple sets of corresponding dialect collection intervals, target semantics and dialect speech to obtain a target dialect recognition model, wherein the target dialect recognition model is used to recognize the target dialect.

2. The method according to claim 1, characterized in that The target dialect area is divided into a plurality of dialect collection intervals from the inner to the outer part with the regional center as the center position, including: Obtaining the postal codes of all regions within the target dialect region to obtain a postal code set, wherein the central postal code of the regional center corresponding to the target dialect region is the smallest postal code in the postal code set, and the multiple postal codes in the postal code set are arranged in ascending order; Using the central postal code as a starting value, the postal codes included in the postal code set are divided into a plurality of equally spaced numerical intervals in ascending order; The regions corresponding to one or more postal codes in the same numerical range are divided into a dialect collection interval to obtain a plurality of dialect collection intervals.

3. The method according to claim 1, characterized in that The step of collecting the dialect speech expressing the target semantics in each of the plurality of dialect collection intervals to obtain a plurality of sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships includes: Sending text data expressing the target semantics to terminal devices located in each of the dialect collection intervals; The voice data returned by the terminal device in response to the text data is received as the dialect voice, and a plurality of sets of dialect collection intervals, target semantics and dialect voices having corresponding relationships are obtained.

4. The method according to claim 1, wherein The step of collecting the dialect speech expressing the target semantics in each of the plurality of dialect collection intervals to obtain a plurality of sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships includes: Identifying speech data expressing the target semantics from the speech database corresponding to each of the dialect collection intervals as the dialect speech; Construct multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships.

5. The method according to claim 1, characterized in that The method of using multiple sets of corresponding dialect collection intervals, target semantics, and dialect speech to train an initial dialect recognition model to obtain a target dialect recognition model includes: Construct the initial dialect recognition model as the initial current dialect recognition model, and repeat the following steps until the target dialect recognition model is obtained: According to the arrangement order of the dialect collection intervals in the target dialect area from inside to outside, the dialect collection intervals are sequentially acquired as current intervals; Training the current dialect recognition model using the target semantics and dialect speech corresponding to the current interval to obtain a first dialect recognition model; If the current interval is not the last interval in the arrangement order, determining the first dialect recognition model as the next current dialect recognition model; In a case where the current interval is the last interval in the arrangement order, the first dialect recognition model is determined as the target dialect recognition model.

6. The method according to claim 1, characterized in that The method of using multiple sets of corresponding dialect collection intervals, target semantics, and dialect speech to train an initial dialect recognition model to obtain a target dialect recognition model includes: Obtaining a target number of initial dialect recognition models, wherein the target number is the number of dialect collection intervals; Using a set of corresponding target semantics and dialect speech to train an initial dialect recognition model, respectively, to obtain the target number of second dialect recognition models; averaging the model parameters of the target number of second dialect recognition models to obtain target model parameters; The initial dialect recognition model having the target model parameters is determined as the target dialect recognition model.

7. The method according to claim 1, characterized in that The regional center of the target dialect region is detected, including: Obtaining the population density distribution of the target dialect area; The area with the highest population density in the target dialect area is determined as the area center.

8. A device for constructing a dialect recognition model, characterized in that: include: a detection module, configured to detect a regional center of a target dialect region, wherein the dialect used in the target dialect region is the target dialect to be identified; a division module, configured to divide the target dialect region into a plurality of dialect collection intervals from the inside out, with the regional center as the center position, wherein the plurality of dialect collection intervals are radially arranged from the inside out, with the regional center as the center, and the closer the dialect collection interval is to the regional center, the higher the similarity between the language used and the target dialect; A collection module, configured to collect dialect speech expressing target semantics in each of the plurality of dialect collection intervals, and obtain multiple sets of dialect collection intervals, target semantics, and dialect speech having corresponding relationships; The training module is used to train an initial dialect recognition model using multiple sets of dialect collection intervals, target semantics and dialect speech with corresponding relationships to obtain a target dialect recognition model, wherein the target dialect recognition model is used to recognize the target dialect.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 7 when executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Recommending method and device based on speech recognition, and medium

    CN110164415A

  • Multi-dialect accent mandarin voice recognition model training method and device, and equipment

    CN112233653A