Differential segmental secure speech transcription

By dividing electronic content into multiple micro-segments and restricting access on distributed computing devices, combined with non-visual viewing tags, the problems of data access and privacy protection in speech recognition systems are solved, improving the security of training data and the integrity of the model.

CN115605947BActive Publication Date: 2026-03-17MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-30
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing speech recognition systems have shortcomings in terms of data access and privacy protection for training data, especially the risk of exposing consumer privacy data during the process of relying on manual labeling, and the company's limited resources make it impossible to effectively train high-quality machine learning models.

Method used

By dividing electronic content into multiple micro-segments and selectively distributing them to multiple distributed computing devices according to security levels, limiting the amount of data accessed by each device, and combining this with a non-visual labeling process, labeled micro-segments are generated to train machine learning models.

Benefits of technology

This improves security in data access and privacy protection of training data, reduces the risk of data exposure, and ensures the integrity and security of training data for machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115605947B_ABST
    Figure CN115605947B_ABST
Patent Text Reader

Abstract

Embodiments are provided for protecting data access to machine learning training data at multiple distributed computing devices. Electronic content comprising raw data corresponding to a preferred data security level is partitioned into multiple micro segments. The multiple micro segments are distributively distributed to multiple computing devices that apply transcription labels to the multiple micro segments. The labeled micro segments are reconstructed into training data that is subsequently used to train a machine learning model while facilitating an increase in data security of the raw data included from the reconstructed micro segments.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Speech recognition technologies used in natural language processing (NLP) rely on large amounts of high-quality labeled data to train deep learning models and other machine learning models for the NLP process.

[0002] One of the most important sources of training data is customer usage data. This data is often matched with specific user scenarios, and NLP models can be very useful in applying them to those specific user scenarios.

[0003] For training data to be effective when training a model, the labels used as the ground truth of the data must be extremely accurate. One more reliable method for obtaining these labels or transcriptions of training data is through human transcription. However, the process of human transcription introduces risks to data security and / or data privacy because it inherently requires a human to listen to customer usage data, including customer voice data, and to provide accurate transcription labels. Therefore, any private information or data, including customer usage data, will be exposed to the human transcriber. Customers retrieving their usage data may not want their private data exposed to human transcribers they do not know, or even those they do know.

[0004] For example, some training data is obtained from consumer speech recognition (SR) devices and systems, where users (i.e., consumers, customers) can use SR devices to input confidential data such as passwords, banking information, and / or credit card information. This data is then sent to a transcription service to label training data used to train and / or fine-tune models used to perform NLP processes. In some examples, data potentially containing private consumer data is retrieved from non-speech data sources such as text, images, and / or video data for use in additional applications (e.g., optical character recognition (OCR) applications).

[0005] Historically, companies in the speech recognition field have relied on third-party vendors to transcribe speech data. In response to the high risks and growing consumer awareness of data privacy, some companies have made changes to their speech transcription practices. For example, in some cases, given that using full-time company employees is safer and less risky than using third-party vendors, some companies now employ full-time staff to perform the tagging of company data. At the very least, companies can use their own secure networks and computers to monitor, track, and regulate the exposure of their data without relying on third-party assurances. However, many companies cannot allocate sufficient resources and staff to transcribe the large amounts of data that need to be transcribed in order to develop enough training data to train machine learning models for speech recognition and other NLP processes.

[0006] Currently, there are no methods for speech recognition systems that have been proven to work more effectively than speech recognition systems trained using manually labeled training data. Therefore, there is a persistent and ongoing need for improved systems, methods, and devices for protecting data access to machine learning training data, including data access to training data obtained from consumer usage data used to train NLP models for performing NLP processing in consumer usage scenarios. Specifically, there is a continuing need and expectation to develop manually labeled process systems and methods to facilitate and / or improve technologies for protecting the privacy and confidentiality of data being labeled.

[0007] The subject matter claimed herein is not limited to embodiments that address any shortcomings or operate only in environments such as those described above. Rather, this background is provided merely to illustrate an exemplary technical field in which some of the embodiments described herein can be practiced. Summary of the Invention

[0008] The embodiments disclosed herein relate to systems, methods, and apparatuses configured to facilitate different levels of data security, and more specifically, to systems, methods, and apparatuses that can be used to protect data access to machine learning training data across multiple distributed computing devices.

[0009] In some embodiments, electronic content is retrieved from a source corresponding to a preferred data security level. The electronic content is divided into multiple micro-segments, wherein the division process is based on the preferred data security level. Once divided, the multiple micro-segments are distributed to multiple computing devices. During distribution, only a certain number of micro-segments from any one source are distributed to the same computing device. In this way, any user(s) without a computing device or any computing device(s) can access the entire electronic content distributed to the computing device. Therefore, the distribution of electronic content in micro-segments can, for example, promote the data security of the underlying data by selectively restricting access to the underlying data.

[0010] The disclosed embodiments include computer-implemented methods for protecting data access to machine learning training data across multiple distributed computing devices. Some disclosed methods include a computing system receiving electronic content containing raw data relating to one or more speakers, wherein the raw data is obtained from the one or more speakers. Once the electronic content is compiled, the computing system determines a security level associated with the electronic content. The computing system then selectively divides the electronic content into multiple micro-segments. Each micro-segment has a duration selected according to the determined security level. After the electronic content is segmented, the computing system identifies multiple destination computing devices configured to apply multiple tags corresponding to the multiple micro-segments. The micro-segments are then selectively distributed to the destination computing devices, while restricting distribution to any particular computing device, such that only a predetermined number of micro-segments from a specific set of raw data will be distributed to any of the destination computing devices.

[0011] In some embodiments, the computing system identifies one or more attributes of a specific differential segment. Where the attributes correspond to an increased level of data security, the computing system further divides the differential segment into fragments.

[0012] In some embodiments, the computing system enables multiple computing devices to apply multiple tags to multiple distributed micro-segments or micro-segment fragments. Once a micro-segment is tagged, the computing system reconstructs the now-tagged micro-segments into reconstructed electronic content, including training data for a machine learning model. Subsequently, the computing system uses the training data from the reconstructed electronic content to train the machine learning model, without exposing the entire electronic content to a single computing device.

[0013] This summary is provided to introduce some concepts in a simplified form, which will be further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter.

[0014] Additional features and advantages will be set forth in the following description, and some features and advantages will be obvious from the description or may be learned by practicing the teachings herein. The features and advantages of the invention can be realized and obtained by the tools and combinations particularly pointed out in the appended claims. The features of the invention will become more apparent from the following description and the appended claims, or may be learned by practicing the invention as set forth below. Attached Figure Description

[0015] To describe how the above-described and other advantages and features can be obtained, a more specific description of the subject matter briefly described above will be presented with reference to specific embodiments shown in the accompanying drawings. It should be understood that these drawings depict only typical embodiments and are therefore not intended to limit the scope. The embodiments will be described and explained with additional features and details using the drawings, in which:

[0016] Figure 1 An example architecture including a computing system is shown, which includes and / or can be used to implement the disclosed embodiments.

[0017] Figure 2 A flowchart is shown as an example method for protecting data access to electronic content used for transcription.

[0018] Figure 3 A flowchart illustrating an example of dividing electronic content into micro-segments is shown.

[0019] Figure 4 A flowchart illustrating an example of removing partial words when dividing electronic content into micro-segments is shown.

[0020] Figure 5 A flowchart is shown as an example method for dividing a differential segment into differential segment fragments.

[0021] Figure 6 A flowchart is shown as an example method for restrictively distributing multiple differential segments to multiple destination computing devices.

[0022] Figure 7 A flowchart is shown for an example method for reconstructing multiple labeled differential segments.

[0023] Figure 8 An example of resolving non-equivalent labels during the refactoring process using voting is shown.

[0024] Figure 9 An example of resolving non-equivalent labels during the refactoring process is shown.

[0025] Figure 10 An example flowchart is shown, illustrating the training of a machine learning model using training data formed from labeled, reconstructed differential segments. Detailed Implementation

[0026] The following discussion relates to methods and method actions that can be performed. While method actions may be discussed in a specific order or shown in a flowchart as occurring in a specific order, this specific order is not required unless specifically stated otherwise, or because an action depends on another action completed prior to its execution. The disclosed embodiments also relate to improved systems, methods, and apparatuses configured to facilitate data security improvements, such as for speech transcription processes and / or the training of machine learning models.

[0027] Figure 1 An example computing architecture 100 for implementing the disclosed embodiments is shown, which includes a computing system 130, which includes various components that will now be described, each of which is stored within a plurality of hardware storage devices 131 and / or one or more remote systems of the computing system 130.

[0028] As shown in the figure, computing system 130 includes hardware storage device(s) 131 and one or more processors 135. Hardware storage device(s) 131 stores one or more computer-executable instructions 131A, 131B, 131... (i.e., code), which are executable by the one or more processors 135 to instantiate other computing system components (e.g., segmenter 134, segment receiver 146, reconstructed content retrieval unit 138, eye-off computing 170, etc.). Processors 135 also execute computer-executable instructions 131A, 131B to implement the disclosed methods, such as... Figures 2-10 The method referenced in [the document / article].

[0029] Although the hardware storage device 131 is currently illustrated as containing computer-executable instructions 131A, 131B, 131... (ellipsis indicating any number of computer-executable instructions), it will be understood that this visual separation from other components included in the computer architecture 100 and / or computing system 130 is merely for the convenience of the discussion herein and should not be construed as representing any actual physical separation of storage components or the distribution of the illustrated components into different physical storage containers. It should also be noted that the hardware storage device 131 may be configured to be segmented (e.g., partitioned) into different storage containers. In some cases, the hardware storage device 131 includes one or more hard disk drives. Additionally or alternatively, the hardware storage device 131 includes flash memory and / or other persistent storage devices and / or combinations of hard disk drives and other persistent storage devices.

[0030] The (multiple) hardware storage devices 131 are currently illustrated as being entirely contained within a single computing system 130. In some embodiments, the (multiple) hardware storage devices are distributed among different computers and / or computing systems, such as (multiple) remote systems (not shown), which are connected via one or more wired and / or wireless network connections. In this case, computing system 130 is considered as a distributed computing system comprising one or more computing systems and / or one or more destination computing devices 160, each computing system and / or destination computing device 160 including its own storage device and corresponding computing components to perform the disclosed functions (such as those described with reference to computing system 130).

[0031] As shown in the figure, the computing system 130 includes a data repository 132 storing initial data 133, such as speech data utterances, audio files, and other initial data (e.g., image data, text data, audiovisual data such as video recordings, or other electronic content). Speech-based data is typically used during speech recognition applications, where the speech data is transcribed (i.e., manually labeled). For example, in some embodiments associated with optical character recognition (OCR) applications, data labeling (i.e., manual labeling and / or annotation) is performed on non-speech and / or image-based data.

[0032] In some cases, data repository 132 also stores corresponding metadata 136 that defines the attributes of initial data 133. In some cases, metadata 136 includes metadata detected from voice service 120, including detected phrase timing, identified text, identification confidence and / or sentence end confidence, as well as other types of metadata (e.g., data type, source, date, format, privacy settings, authorship or ownership, etc.). In some cases, data repository 132 also includes tagged data 168, which will be described in more detail below. Additionally, although not shown, in some cases, data repository 132 also includes or includes hardware storage devices 131 storing computer-executable instructions for implementing the functions disclosed herein.

[0033] In some cases, the data stored in the data repository 132 is encrypted, such as the initial data 133 being encrypted, and the initial data is made directly accessible to no person or computing entity without restriction and / or with eyes-on access.

[0034] Although data repository 132 is currently shown as distributed, it will be understood that data repository 132 may also include a single database and may not be distributed across different physical storage devices.

[0035] It should also be understood that data repository 132 stores electronic content from any combination of one or more different data sources 110. For example, the electronic content of initial data 133 (and corresponding metadata 136) may include voice or audio data used for and / or from one or more different speech recognition systems and / or voice services 120. Electronic content received from data source 110 may be discrete content files and / or streaming content.

[0036] Segmentation divider 134 divides the complete utterance into micro-segments (i.e., slices, fragments, segments, etc.) based on strategy 140 (i.e., segmentation, partitioning, parsing, truncation, interruption, etc.). Segmentation divider 134 also supervises the processing of overlapping content of micro-segments and monitors for any errors in speech segmentation introduced by any process implemented by any component of the computer architecture 100. In some cases, segmentation divider 134 receives electronic content from data repository 132, wherein the electronic content includes sentences (i.e., utterances and / or candidates for segmentation) that are prioritized for human review based on identification confidence and sentence confidence scores associated with the electronic content. In some embodiments, these identification confidence and sentence confidence scores, including other attributes, are included in the metadata corresponding to the electronic content stored in data repository 132 storing initial data 133 and / or stored in data repository 132 storing metadata 136.

[0037] In some embodiments, the segment divider 134 is configured to implement any of the methods for segmentation, such as... Figures 2-7 As shown in the diagram. In some embodiments, segmenter 134 divides electronic data into microsegments on demand (e.g., in response to a request for segmentation and / or a request for transcription of electronic content). Additionally or alternatively, segmenter 134 segments / divides electronic content before a request is made, for example, when electronic content (with or without associated metadata) is identified as being stored in data repository 132, and when such data is available for segmentation.

[0038] In some embodiments, the segmentation divider 134 is configured to include multiple modes (e.g., a simple mode and / or a smart mode). In some cases, when the segmentation divider 134 operates in simple mode, the segmentation process implemented by the segmentation divider 134 is associated with a static segmentation process. In some embodiments of static segmentation, each utterance is divided into micro-segments of equal duration and / or each utterance is divided into micro-segments based on a global policy 140, which in some cases is static and / or at least invariant during segmentation.

[0039] In some cases, when the segmenter 134 operates in intelligent mode, it dynamically divides the electronic content into micro-segments using sentence-end confidence scores or other metadata that can be used to identify natural speech interruptions and / or associated data security levels. Therefore, the segmenter 134 can be implemented as a machine learning model, a deep learning algorithm, and / or any machine learning model, trained and / or fine-tuned to learn the optimal pattern and corresponding optimal strategy for the segmentation process based on the attributes of the electronic content.

[0040] In some embodiments, the segment divider 134 performs dynamic segmentation when operating in intelligent mode, including dividing the utterance into micro-segments of varying durations based on the unique characteristics or attributes of the utterance identified during the segmentation process. In this way, during the segmentation process, different strategies (such as strategy 140 from multiple available strategies) are selected and uniquely applied to each utterance and / or micro-segment based on the detected attributes of the segmented utterance.

[0041] In some embodiments, the segment divider 134 is configured to segment non-speech data according to several methods disclosed herein, wherein text in the OCR image is divided into micro-segments (i.e., the original image is subdivided into multiple smaller images). In this way, the smaller images (i.e., image segments, fragments, frames, slices, etc.) include only a portion of the original image, or in other words, the smaller images include only a few words from the multiple words included in the complete original image.

[0042] It should be understood that one or more strategies 140 are configured to determine the rules or guidelines for the segmentation process implemented by the segmentation divider 134. In some embodiments, strategy 140 includes determining media time (i.e., duration). For example, for consumer data (consumer data corresponding to a data security level), strategy 140 specifies a predetermined time window or duration (e.g., 8 seconds) that can be mapped to multiple words (e.g., 7-10 words). For more sensitive data, or data corresponding to an increased data security level compared to consumer data, strategy 140 specifies a shorter window (e.g., less than, for example, 2-3 seconds) to protect the content of data retrieved from data repository 132 (e.g., initial data 133).

[0043] In some embodiments, strategy 140 includes a predetermined number of words (e.g., a limit of N words). In this case, word boundaries of the electronic content are determined for individual words appearing in the electronic content retrieved from data repository 132 (e.g., initial data 133). In some cases, metadata 136 corresponding to the electronic content includes word counts and word boundary information, which are then used to determine the segmentation of the limit of N words.

[0044] Alternatively or concurrently, one or more strategies 140 are based on identified text. For example, segmentation 134 can be configured to break after N consecutive numbers or characters. This is beneficial for electronic data or subsequent microsegments of electronic data that include number-based keywords (i.e., number-based attributes), such as credit card numbers, identity numbers (e.g., social security, driver's license, and / or passport numbers), transaction confirmation numbers, bank account information, routing numbers, passwords, and / or other attributes identified by computing system 130, wherein the attributes correspond to a preferred data security level associated with the electronic content.

[0045] In some cases, attributes and / or identifier keywords (attributes and / or keywords corresponding to the data security level, which are based on numeric or non-numeric values) are segmented between micro-segments or between micro-segment fragments (i.e., micro-segments further divided into segments). In some cases, attributes and / or keywords are omitted from the micro-segments.

[0046] In some embodiments, strategy 140 further includes segmentation rules based on speech prediction of silence and / or based on the probability of micro-segments containing speech. In some embodiments, strategy 140 instructs segmenter 134 to segment electronic content based on natural speech interruptions identified in the electronic content. When using natural speech interruptions, subsequent micro-segmentation advantageously reduces the probability of including partial or truncated words.

[0047] Therefore, in some embodiments, strategy 140 refers to the static determination of the segmentation process and method, wherein a predetermined length and / or duration of each micro-segment is used to divide the electronic content. In some embodiments, the predetermined length corresponds to the identified attributes of known data sources from a plurality of data sources 110 and / or data sources from which the electronic content is obtained. In some embodiments, strategy 140 results in dynamic segmentation (i.e., segmentation is implemented using different applied strategies based on the attributes of the segmented electronic content).

[0048] In some embodiments, the attributes of the electronic content (e.g., initial data 133) and the attributes of the subsequently identified micro-segments are stored in a data repository 132 containing metadata 136. After the authorization process, this metadata 136, together with the micro-segments generated by the segmenter 134, is retrieved and / or sent (i.e., a push and / or pull data system) to the segment receiver 166 of the computing device 166.

[0049] The segment receiver 166 ensures that the same client (i.e., computing device) does not acquire more than X number of micro-segments, and / or that at a given time, Y% of a given discourse (i.e., a set or subset of electronic content) is not retrieved for tagging.

[0050] Conceptually, the segmentation process, also referred to in this paper as the micro-segmentation process associated with forming micro-segments, is similar to dividing a credit card into multiple bins and ensuring that the contents of these bins are not available to any single customer in order to enhance the security of the data associated with the credit card.

[0051] Additionally, in some embodiments, requests for tagged micro-segments are protected and verified for user access. For each initial data item 133 that is not entirely manually tagged, a maintained hash list exists that maps to a hash of the user's percentage segmentation list length, and this maintained hash list is associated with the destination device requesting / receiving different data. In some embodiments, multiple transcriptions exist for a single initial data item 133 (i.e., electronic content). The disclosed system maintenance maps hashes to trackers / tracking information of the recipient / requesting system until the transcription of the complete utterance or the target portion of the initial data 133 is completed. In some embodiments, segmentation is performed via partial encryption and keychain application.

[0052] This configuration beneficially facilitates the computing system's ability to track every request from each computing device, thereby ensuring that the computing system will satisfy the policy 140 associated with the differential segment. Furthermore, a stored database exists (even after the differential segment has been reconstructed) that stores which differential segments were transcribed by which computing devices.

[0053] In some embodiments, before computing device 160 can access the microsegments of electronic content and associated metadata, the identifier of the computing device must be confirmed and authorized as a computing device with permission to access the microsegments. In some examples, the computing device corresponds to an identifier among a plurality of identifiers indicating a previously authorized computing device. In some cases, authorization is based on a predetermined threshold corresponding to a data security level associated with the electronic content, wherein based on the data security level associated with the content and / or based on security permissions associated with different recipient systems, some recipient computing devices are allowed to access and / or send a defined set of data content reaching a certain absolute amount and / or percentage of the total set of data segments.

[0054] In some embodiments, the authorization process will be predetermined and / or pre-configured. In some embodiments, whenever a computing device requests access to one or more microsegments and / or a device is identified as a potential target for receiving microsegments, the device is screened for authorization and authorized via an API authorization platform before the microsegments(s) are sent to the device. It should also be understood that the identifiers corresponding to the computing device and / or microsegments and / or microsegment metadata are encrypted via an encryption system.164

[0055] Once a computing device is authorized and receives one or more microsegments, it is able to apply one or more tags to the microsegments via visually inspected human marking 180. In some embodiments, data tags 168 are sent to a data repository 132 storing tagged data 150, the data storage device 132 including the distributed microsegments, associated metadata, and corresponding transcription tags. In some embodiments, the tagged data 150 is reconstructed electronic content, wherein the microsegments have been reconstructed based on the original ordering of the microsegment content before being partitioned from the original electronic content. It should be understood that data tags 168 may be referred to as transcription tags and / or transcription and / or simply as the tags described herein and / or referenced throughout the accompanying drawings.

[0056] Once the data segments are tagged, the reconstructed content retrieval unit 138 retrieves a reconstructed set of multiple segments of the tagged data 150, metadata 136, and / or other initial data 133, where the computation system performs non-visualized computation 170 on the complete session data. This computation includes training of models, algorithms, and / or other data post-processing.

[0057] By leveraging non-visualized computation, segmentation, and reconstruction, the computational system and the overall transcription process enhance the security of underlying data and reduce privacy risks (i.e., improve data security) in end-to-end systems. Not only is data exposure minimized during transcription—the generation and use of training data for machine learning models—but the same principle can be applied during the evaluation, testing, and / or debugging of machine learning models. In some cases, model training is controlled within compliance boundaries, and some data metrics, such as word error detection, can be retrieved from these boundaries. In other cases, evaluation and / or debugging entities may only have access to micro-segments containing a large number of errors.

[0058] In this way, visually inspected tags are restricted to accessing only specific portions of the electronic content from which they are parsed, while human entities or computing devices corresponding to human entities are not authorized to access the full electronic content, thereby maintaining the data security level of the initial data 133, which may include sensitive consumer data.

[0059] In some embodiments, data access is controlled by user permissions and follows compliance rules, such as any user activity and audit log activity associated with the JIT (Just in Tim) application. Through the data access service, the computing device (and / or the corresponding user) can access encrypted microsegments of the initial data 133. In some embodiments, the encryption is an AES-128 key or even more secure. The voice data is expected to be encoded as MP4 in MPEG-DASH format and is not downloadable. Encryption keys and playback asset links generated for microsegments or subsets of microsegments for the computing device (and / or the transcriber) will expire within policy-based limits, such as five minutes.

[0060] Methods for generating differential segments

[0061] Now let's turn our attention to... Figure 2 It shows a flowchart 200, which includes components related to those referenced above. Figure 1 The computing system 130 described herein can implement various actions associated with exemplary methods, and / or various actions associated with protecting data access to machine learning training data during the labeling process of corresponding data.

[0062] like Figure 2 As shown, flowchart 200 and the corresponding method include the action of a computing system receiving electronic content comprising raw data (action 210). The computing system then determines a security level associated with the electronic content (action 220). Subsequently, the computing system selectively divides the electronic content into multiple micro-segments according to duration, the duration of which is selected at least in part based on the determined security level (action 230). Subsequently, the computing system identifies multiple destination computing devices configured to apply multiple tags corresponding to the multiple micro-segments (action 240). Finally, the computing system selectively and restrictively distributes the multiple micro-segments to the multiple destination computing devices (action 250). It is worth noting that the action of distributing micro-segments (action 250) in some cases also includes restricting or otherwise constraining the distribution of micro-segments, such that only a predetermined number of micro-segments from the raw data (or, a predefined set of data from the raw data) are distributed to any of the destination computing devices.

[0063] It should be understood that in some cases, electronic content (action 210) includes initial data 133 stored in data repository 132 and / or initial data 133 obtained directly from one or more remote data sources 110 and / or (multiple) voice services 120.

[0064] Additionally, in some cases, the raw data (or a predefined set of raw data) includes consumer usage data, which includes private, sensitive, and / or confidential data from one or more consumers. Therefore, the security level determined for electronic content identification (action 220) will be based on attributes of data source 110, such as whether it is a private or public data source and / or whether the speakers associated with the data source have pre-authorized sharing the data collected by the voice service. In some cases, attributes are identified by the data source / voice service. In other cases, attributes are independently and automatically identified by computing system 130 during the processing of electronic content when it is received. In some cases, attributes are specified using metadata associated with the electronic content when the electronic content is received and / or after the electronic content is processed.

[0065] The action of dividing the electronic content into micro-segments will be implemented by the segment divider 134, such as... Figure 1 As shown in the image. The process of dividing electronic content, also known as segmenting, cutting, or dividing electronic content into micro-segments, will be described in more detail below.

[0066] Refer to the action of multiple destination computing devices (Action 240), the destination computing devices can be, for example Figure 1 The (multiple) computing devices 160 represent one or more computing devices. This action can be performed by the computing system identifying the destination computing device based on explicit user input specifying multiple destination computing devices, and / or can be based on various other factors and rules (such as matching the attributes of the electronic content with the specific attributes of the destination computing device selected from more multiple systems / devices), based on routing protocols, based on load balancing rules, etc.

[0067] Note that the multiple destination computing devices (Action 240) are distributed and / or independent computing devices, where one destination computing device does not communicate with another computing device and / or cannot access or is unaware of the micro-segments (and / or may be represented as corresponding tags) received by another computing device. In some embodiments, the destination computing device corresponds to a transcription entity, wherein the entity is an artificial transcriber.

[0068] It should also be noted that the action (action 250) used for restrictively distributing multiple differential segments will be handled by... Figure 1This is implemented using one or more of the various components of the computer architecture 100, such as computing device 160, API authorization and authentication 162, encryption 164, segmented receiver 166, human-tagged data viewing 180, and / or data repository 132 storing metadata 136. The rules for selectively distributing electronic content to destination computing devices will be described in more detail below. Note that in some embodiments, the term "human-tagged" as used herein refers to the ability of a human observer or commentator to access a collection or subset of data. Furthermore, the term "non-human-tagged" as used herein is used to describe data access limited to internal computing systems (i.e., non-human entities).

[0069] Now turn attention to Figures 3-4 , Figures 3-4 Exemplary methods and examples are shown that are associated with dividing electronic content into micro-segments (e.g., separating, segmenting, splitting, cutting, or otherwise segmenting) for use in tagging electronic content.

[0070] In some embodiments, such as Figure 3 As shown, the electronic content includes the original utterance 310 [The dogs ran around and around the firepit in the backyard], which can be used for segmentation. A security level corresponding to the original utterance 310 is determined. Once the security level is determined, a specific strategy is selected from multiple strategies (i.e., Figure 1 Strategy 140), where this particular strategy specifies that the micro-segments will include a maximum number of words, such as no more than six words. In this example, the original utterance 310 is then divided into multiple micro-segments, specifically two micro-segments of the divided utterance 320 (i.e., [The dogs ran around and around] and [the firepit in the backyard]), where each micro-segment includes up to six words. This is an example of static segmentation.

[0071] In some embodiments, two microsegments comprising a first segmented utterance 320 are distributed to multiple destination computing devices. Additionally or alternatively, to facilitate diverse distribution of a particular microsegment, segmentation can be performed to create overlap between segments, which can facilitate labeling different segments with different contexts to promote more accurate labeling and / or verification of label accuracy. For example, in one example, the original utterance 310 is further segmented into a second segmented utterance 330 (or, a second set of microsegments) comprising a set of three microsegments, including (1) [The dogs ran], (2) [around and around], and (3) [the firepit in the backyard]. In this case, the original utterance 310 is segmented based on natural speech segmentation, for example, where slight pauses are detected between “ran” and “around” and between “around” and “the”. The initial segmentation is an example of dynamic segmentation, generated by the identification of pause attributes during segmentation.

[0072] In this example, the second set of microsegments (divided discourse 330) reflects the overlap between natural breaks in parts or words of the microsegments of the divided discourse 320, which can be used to provide context in the microsegments for statistically more accurate transcriptional markers.

[0073] In some embodiments, a segmentation process that divides a discourse into two or more sets of segments (e.g., segmented discourse 320 and segmented discourse 330, or segmented discourse 320 and segmented discourse 340, etc.) facilitates the generation of multiple tags for a discourse or a portion thereof. In some embodiments, multiple tags are generated for portions of a discourse (e.g., discourse 310) by dividing the discourse into multiple segments, wherein each subsequent segment includes a portion of a previous segment (e.g., segmented discourse 360).

[0074] Additional or alternative, based on changing strategies (e.g., Figure 1 Strategy 140), the original utterance 310 can also be divided into third and / or fourth sets of micro-segments (see, divided utterances 340, 350). Note that, based on words, durations, or other (multiple) features, the original utterance 310 can be divided into any number of sets of micro-segments and / or have any number or configuration of overlapping micro-segment portions. Each set of micro-segments can be divided based on different attributes identified in the original utterance.

[0075] In some examples, segmented utterance 320 indicates a first security level corresponding to the original utterance 310. Subsequently, segmented utterance 350 indicates a second security level corresponding to the original utterance 310, wherein the micro-segments of segmented utterance 350 comprise 2 to 4 words, indicating a higher security level than the first security level. In some embodiments where the second security level corresponds to segmented utterance 320, the attributes determined in the micro-segments of segmented utterance 320 are further segmented. Therefore, in some cases, the parsed phrase of segmented utterance 350 is a micro-segment of the original utterance 310 and / or a micro-segment fragment of the micro-segment of segmented utterance 320. These embodiments are particularly useful when attempting to add security to underlying data such as cryptographic data, financial transaction data, and personally identifiable information.

[0076] It should be understood that the original discourse 310 can be divided into divided discourse 320, divided discourse 330, divided discourse 340, divided discourse 350, divided discourse 360 ​​and / or any combination of these and other divided discourses, for transmission to the receiving party's marking device / system. See also the following in this document. Figure 5 Description of the exemplary method shown.

[0077] Now let's turn our attention to... Figure 4 A flowchart illustrating an example of segmenting a discourse from a raw discourse. In some embodiments, the segmentation process may produce partial or truncated words that are difficult to transcribe accurately. In such cases, partial words are removed from the microsegments before they are distributed to a computing device for labeling. For example, raw discourse 410 is shown as [The dogs ran around and around the firepit in the backyard]. Various methods are used to identify word boundaries, including based on word timing and the predicted beginnings and ends of words. Once word boundaries are identified and defined (see discourse 420), raw discourse 410 is divided into multiple microsegments, such as segmented discourse 430. In some embodiments, word boundaries are identified by buffering at both ends to ensure that key beginnings and / or ends of words are not missed.

[0078] In the case shown and provided only as an example, the original utterance 410 is divided into differential segments with a fixed duration of five seconds (i.e., a static segmentation strategy). However, the first and second differential segments include local words (e.g., "around" is segmented into "a" and "round"). The computational system (e.g., Figure 1The computational system 130 identifies local words by comparing the word boundaries of the original utterance 410 with the word boundaries identified in the segmented utterance 430. In this case, “a” and “round” are identified as local words and removed from the segmented utterance (see, modified segmented utterance 440). In some embodiments, the identification and / or removal of local words triggers the computational system (e.g., Figure 1 The computational system 130 generates another set of differential segments that do not contain any local words and still satisfy the initial requirement of a duration of less than or equal to five seconds (see, partitioned discourse 450).

[0079] In some embodiments, after the electronic content is divided into micro-segments, the micro-segments are further modified before being distributed to multiple computing devices. In some embodiments, the micro-segments are further modified after distribution and tagging, but before being reconstructed into reconstructed electronic content. In some embodiments, the reconstructed, labeled micro-segments are further modified before being used as training data for a machine learning model.

[0080] In any of the foregoing embodiments with modifications, the modifications may include truncating the microsegment at a natural speech interruption identified by the computing system, removing local words identified by the computing system in the microsegment, changing the pitch of the speech audio included in the microsegment, slowing down or speeding up the frequency of the speech audio included in the microsegment, increasing room inresponse noise perturbation, adding white noise or other background noise, and / or making other sound or audio modifications.

[0081] It is worth noting that altering the pitch of speech audio enhances data security because, in cases of transcribed entity recognition or memorizing specific voice signatures, changing the pitch renders the speech audio unrecognizable by computing devices and / or human transcribers. Furthermore, in some embodiments where specific microsegments are distributed multiple times to a single computing device, the initial distribution of the microsegment to the computing device comprises the original speech audio, and subsequent distributions of the microsegment to the same computing device comprise modified speech audio (e.g., by changing the pitch).

[0082] In some embodiments, the computing system selects different strategies (e.g., such as...) Figure 1Strategy 140 (as shown in the strategy) is used to generate alternative and / or additional sets of differential segments. For example, to generate the segmented utterance 420, differential segments equal to a predetermined duration (e.g., five seconds) are generated. In some cases, the segmented utterance 450 includes generated differential segments equal to or less than three seconds. Alternatively, the segmented utterance 450 includes generated differential segments comprising equal to or less than N words (e.g., no more than three words). The additional and / or alternative segmented utterances 450 are provided to a computing device for tagging to ensure that all words in the original utterance 410 receive at least one corresponding transcription tag.

[0083] Now let's turn our attention to... Figure 5 It shows a flowchart 500, which includes components related to those referenced above. Figure 1 The computing system 130 described can implement various actions associated with exemplary methods, and / or various actions associated with protecting data access to machine learning training data during the labeling process of corresponding data.

[0084] like Figure 5 As shown in the flowchart 500 and the corresponding method, the computing system receives electronic content including raw data (action 510). The computing system then determines a security level associated with the electronic content (action 520). Thereafter, the computing system selectively divides the electronic content into multiple micro-segments according to their duration, the duration of which is selected based on the determined security level (action 530). Subsequently, the computing system identifies one or more keywords associated with a second security level (action 540). The second security level is an increased security level compared to the initially determined security level associated with the electronic content. After identifying the keywords in the micro-segments, corresponding to the increased security level, the computing system divides the micro-segments (multiple) into one or more micro-segment fragments (action 550).

[0085] In some embodiments, one or more keywords (i.e., attributes of electronic content and / or microsegments) include one or more of the following: personal identification numbers (driver's license numbers, passport numbers, social security numbers, etc.), transaction data (credit card and / or debit card numbers, security codes, personal identification numbers (PINs), transaction confirmations, remittances, currency amounts, bank accounts, routing numbers, cheque numbers, etc.), password data, recovery emails, telephone numbers, keyword phrases or words, other series of characters and special characters (ASCII and / or non-ASCII characters), number sequences, one or more names, one or more characters identified as having abnormal speech, and / or terms indicating data associated with user or account credentials. These types of data may be referred to as credential data, authentication data, or verification data.

[0086] In some embodiments, keywords are identified by prior trigger words. For example, if a credit card appears at the beginning of a utterance and / or in the first utterance, a series of numbers appearing at the end of the utterance and / or in the second utterance are marked as keywords corresponding to an increased security level. In this case, the first utterance is divided into multiple micro-segments, and the second utterance is divided into multiple micro-segments and further into micro-segment fragments, such that a series of numbers is parsed between the micro-segment fragments. In some embodiments, only keywords are divided into fragments. Additionally or alternatively, micro-segments comprising only one or more keywords are divided into micro-segment fragments. In some embodiments, where one or more keywords are identified in at least one micro-segment fragment, each micro-segment of the collection or subset of electronic content is subdivided into micro-segment fragments.

[0087] In some embodiments, micro-segment segments are formed by splitting or otherwise dividing keywords into fragments or local keywords. Additionally or alternatively, micro-segment segments are generated by removing one or more keywords from the micro-segments, and / or by removing one or more keywords from the electronic content and then dividing the electronic content into micro-segments.

[0088] In some embodiments, the determined security level (action 520) is a first security level associated with the electronic content, wherein the computing system operates on the electronic content to divide it into multiple micro-segments. Therefore, the identification of keywords corresponding to an increased security level (which is a second security level) causes the computing system to operate on the micro-segments to further divide them into micro-segment fragments.

[0089] In some embodiments, after determining a first security level, the computing system identifies one or more keywords corresponding to the increased security level in the electronic content and automatically divides the electronic content into micro-segment fragments.

[0090] Regarding the security level, it should be understood that the security level may be based on the attributes of the audio discovered by the computing system, and / or on the metadata received with or associated with the audio when the audio is received from a third party (e.g., an audio file with a metadata descriptor), and / or may be obtained by looking up attributes in a database or table storing attributes of different electronic content (including the audio being processed).

[0091] In some embodiments, keywords are identified by a segmentation divider, for example... Figure 1 The segmentation divider 134. In some embodiments, keywords are identified by a voice service, such as... Figure 1One or more voice services 120, wherein information about keywords is included in metadata about electronic content stored in a data storage device, such as a data repository 132 storing initial data 133 and / or storage. Figure 1 The metadata of 136 is stored in the data repository 132.

[0092] Segmentation of speech data is highly beneficial for several reasons. In cases where long data sessions need to be transcribed, most traditional speech transcription services and / or NLP machine learning models are unable to handle large amounts of data (i.e., in terms of storage capacity versus available processing memory and / or word length and / or duration length and / or the number of speakers contributing to the data).

[0093] In some embodiments, the transcription process is implemented by manually reviewing the conversation within a permissible length. However, any conversation exceeding this limit is not reviewed. In some embodiments, the transcription process includes dividing the conversation into segments of permissible length. Time-based microsegmentation may cut audio in the middle of words and may sometimes fail to provide sufficient context for accurate manual review. Therefore, in some embodiments, the disclosed segmentation process advantageously includes sentence-end prediction and permissible length time constraints to perform segmentation. For example, “this is a great novel, and I highly recommend it for everyone to read” can be segmented into “this is a great novel,” “and,” “I highly recommend it,” and “for everyone to read”—where each microsegment is divided in such a way that its duration / length is below a predetermined threshold length and / or controlled by natural speech interruptions.

[0094] Furthermore, segmenting data into micro-segments and restricting the distribution of these micro-segments helps improve data security because no single computing device and / or transcriber has sufficient context to decipher the meaning of the data segmented into micro-segments.

[0095] Selective and / or restrictive distribution of differential segments

[0096] Now let's turn our attention to... Figure 6 It shows a flowchart 600, which includes components related to those described in the above reference. Figure 1The computing system 130 described can implement various actions associated with exemplary methods, and / or various actions associated with protecting data access to machine learning training data during the labeling process of corresponding data. Specifically, in the process of restrictively distributing electronic data to multiple distributed computing devices, the content in the microsegments is labeled by human and / or automatic systems based on the observable features / attributes of the microsegments.

[0097] like Figure 6 As shown, flowchart 600 and the corresponding method include the action (action 610) of a computing system determining a certain number of micro-segments from electronic content that have been distributed to a first destination computing device. The first destination computing device is one of a plurality of distributed computing devices. The computing system then identifies a micro-segment to be distributed to any of the plurality of destination computing devices (action 620). After identifying a specific micro-segment, the computing system determines that distributing that micro-segment to the first destination computing device would cause the total number of micro-segments distributed to the first destination computing device to exceed a predetermined threshold (action 630). Subsequently, the computing system determines a certain number of micro-segments from electronic content that have been distributed to a second destination computing device (action 640). The second destination computing device is one of a plurality of distributed computing devices, for example... Figure 1 The computing device 160. Before distributing the identified differential segment to the second destination computing device, the computing system determines that the distribution of the identified differential segment will not cause the total number of differential segments distributed to the second computing device to exceed a predetermined threshold (action 650). When it is determined that the threshold will not be exceeded, the computing system then distributes the differential segment to the second computing device (action 660) and updates the records to reflect the total distribution of the differential segment(s) to the second computing device.

[0098] In some embodiments, a predetermined threshold is determined based on the security level associated with the electronic content. In some embodiments, the threshold is an upper limit on the total maximum value of the differential segments. In some embodiments, the threshold is an upper limit on the total maximum value of the differential segments of a specific set or subset of electronic content from a shared source.

[0099] In some embodiments, based on a determined security level (see...) Figure 5 Action 510) divides the micro-segments from the given number of micro-segments (Action 610) and / or the identified micro-segments (Action 620). In some embodiments, the micro-segments referenced in flowchart 600 are divided randomly or based on multiple different characteristics (different from and / or including a determined security level). Therefore, in some embodiments, the restricted distribution of micro-segments does not depend on micro-segments that have already been divided based on a determined security level associated with the electronic content from which the micro-segments are obtained.

[0100] Additionally, or optionally, the restricted distribution of microsegments is performed based on a determined security level of the microsegments (and / or the electronic content from which the microsegments are obtained). For example, if a particular microsegment or set of microsegments is determined to have a low security level (below a predetermined threshold), the microsegments are freely and / or randomly (i.e., non-restrictively) distributed to one or more computing devices. In some embodiments, the restricted distribution of these microsegments is triggered (i.e., activated) by the determination of a specific security level of one or more microsegments and / or the electronic content from which the microsegments are derived. Furthermore, in some cases, one or more thresholds (actions 630, 650) are fine-tuned and / or adjusted based on the determined security level of the microsegments (and / or the corresponding electronic content).

[0101] Therefore, it should be understood that in some cases, one or more actions corresponding to the methods associated with flowchart 400 (and / or flowchart 500) and one or more actions corresponding to the methods associated with flowchart 600 are performed independently of each other. However, in some cases, the execution of one or more actions associated with flowchart 600 depends on the execution of one or more actions associated with flowchart 400 (and / or flowchart 500).

[0102] In some embodiments, the disclosed method further includes the action of the computing system accessing a lookup table including identifiers, wherein each identifier corresponds to one of a plurality of destination computing devices. The computing system then determines a certain number of micro-segments for each destination that has been distributed to the destination computing devices. Thereafter, the computing system links the identifiers to that certain number of micro-segments for each destination device that has been distributed to the destination computing devices, wherein the certain number of micro-segments corresponds to a percentage of the total number of micro-segments distributed to each destination device from the electronic content.

[0103] In some embodiments, the disclosed method further includes a computing system determining whether a particular microsegment to be distributed to a second or receiving computing system is contiguous with a previous microsegment (from a predetermined set of partitioned data) that has already been distributed to the receiving computing system from a specifically defined data set. Then, if it is determined that one microsegment from a set of microsegments that created the same underlying data set is contiguous with another microsegment, the system restricts the distribution of that particular microsegment to the receiving computing system.

[0104] In some embodiments, the computing system also determines (before specific distribution of a specific microsegment) that distributing the specific segment to a second destination computing device will not cause the total number of microsegments distributed to the second destination computing device to exceed a predetermined threshold, while still facilitating the acquisition of multiple transcription tags for the specific microsegment.

[0105] It should be understood that microsegments will be distributed according to various methods, where the selectivity or restriction of distribution is based on various criteria. For example, in some embodiments, the microsegment threshold for each computing device is the maximum value or number of microsegments. In some embodiments, microsegment distribution is limited based on a threshold for how many consecutive microsegments a computing device receives. In this case, there may be no upper limit to the total number of microsegments as long as the computing system does not receive X number of consecutive microsegments. In some embodiments, the threshold is fixed. In some embodiments, the threshold is based on a determined security level of the electronic content and / or on keywords identified by the computing system that are associated with an increased data security level.

[0106] Conflict solutions for reconstructed electronic content

[0107] Now let's turn our attention to... Figure 7 It shows a flowchart 700, which includes components such as those referenced above. Figure 1 The computing system 130 described herein can implement various actions associated with exemplary methods, and / or various actions associated with protecting data access to machine learning training data during the labeling process of corresponding data, including techniques for resolving conflicts between non-equivalent labels for specific differential segments.

[0108] In some embodiments, after the electronic content is divided into micro-segments, the micro-segments are distributed to multiple computing devices configured to apply transcription tags to the micro-segments (e.g., see...). Figure 1 (Data tag 168). In some cases, the same microsegment will be distributed to more than one computing device, where different computing devices will generate copies or multiple transcription tags for the same microsegment. However, sometimes, two or more transcriptions of a particular microsegment may be inconsistent, or the same word or part of the microsegment may have different tags.

[0109] In some embodiments, electronic content is divided into two or more sets of micro-segments, wherein one or more micro-segments (or portions of micro-segments) of a first set overlap with one or more micro-segments (or portions of micro-segments) between the two sets. Therefore, when these sets of micro-segments are distributed to different computing devices, the tags returned from the computing devices for the overlapping portions are sometimes different. See, for example, [link to relevant documentation]. Figure 3 and Figure 4 This paper describes several methods for performing conflict resolution to determine which of multiple transcription tags will be included in the generation of training data for machine learning models configured for speech recognition or NLP applications.

[0110] like Figure 7 As shown, flowchart 700 and the corresponding method include the action of a computing system reconstructing multiple micro-segments into reconstructed electronic content (action 710). The computing system then receives multiple tags corresponding to the multiple micro-segments divided from the electronic content (action 720). Thereafter, the computing system determines that at least one overlapping portion of the micro-segments includes a set of non-equivalent corresponding tags (action 730). After determining the non-equivalent repeating tags for the micro-segments, the computing system includes a specific tag from the set of non-equivalent corresponding tags in the reconstructed electronic content for the overlapping portion of the micro-segments (action 740).

[0111] In some embodiments, before the microsegments are distributed for transcriptional tagging, the microsegments are advantageously reconstructed or reordered to reflect the initial partitioned discourse (the partitioned discourse comprises multiple microsegments partitioned from electronic content, e.g., see...). Figure 3 The utterances (320, 330, 340, 350) are divided into temporal and / or sequential orders to maintain the data integrity of the electronic content and are used to generate training data. Additionally or alternatively, the micro-segments and corresponding labels can be reconstructed based on matching security levels, similar keywords, matching and / or similar data sources, similar speaker attributes, or other attributes identified in the micro-segments.

[0112] In some embodiments, only one transcription tag is selected to be included in the reconstructed electronic content used for training data. In some embodiments, one or more transcription tags are selected to be included in the reconstructed electronic content used for training data. In some cases, multiple transcription tags are included as weights for each transcription tag when evaluating the efficiency and / or accuracy of a machine learning model trained with the training data. It should be understood that the reconstructed electronic content can be used for training and / or evaluation of machine learning models.

[0113] Now let's turn our attention to... Figure 8 and Figure 9 This illustrates an example of parsing non-equivalent transcription tags for specific microsegments. For example, Figure 8 This shows examples including those from references such as those mentioned above. Figure 1 The exemplary methods that the computational system 130 described can implement are associated with, and / or examples of various configurations associated with protecting data access to machine learning training data during the labeling process of the corresponding data. More specifically, these methods are associated with resolving conflicts between non-equivalent labels for a specific differential segment.

[0114] Figure 8 A utterance 810 is shown, comprising the audio content of a specific speaker uttering “Yesterday two dogs ran around and around.” Audio frequencies 820 recorded from and associated with utterance 810 are shown below utterance 810. As shown in lines 840, 850, and 860, audio recording 820 is considered to be divided into multiple micro-segments 830 and distributed (with limited distribution) to different destination systems for labeling. The portions of the audio recording utterance 820 are shown as letter labels a, b, c, d, e, and f, where each micro-segment comprises one or more audio recording portions (it should be understood that micro-segments can be generated and distributed via any of the multiple partitioning and distribution methods described herein).

[0115] In this example, the set of differential segments is distributed to multiple receiving computing devices (A, B, C), each of which is configured as follows: Figure 1 The computing device 160. In the example shown, different microsegments are created and distributed to computing device A (i.e., microsegments [a, b, c] and [d, e, f]), to computing device B (i.e., microsegments [a], [b, c, d] and [e, f]), and to computing device C (i.e., microsegments [a, b], [c, d] and [e, f]), such that each of them obtains a complete set of electronic content for discourse 810 (a, b, c, d, e, and f). This example is provided to illustrate how conflict resolution occurs, while recognizing that in some cases, and as previously stated, based on security strategies (where applicable), microsegments of a complete discourse or consecutive microsegments of a discourse will not be sent to any particular computing system. This example also illustrates how different micro-segments of the same utterance can be distributed to different devices for labeling, where in some cases, the different micro-segments will overlap with different parts of the underlying electronic content (e.g., as shown in box 870, where a micro-segment (b, c, d) of the content sent to device B will overlap with micro-segments (a, b) and (c, d, e) of the content sent to device C).

[0116] exist Figure 8 In the figure, the reconstructed micro-segments include reconstructed or tagged electronic content 842, 852, and 862, which correspond to tags provided by computing devices A, B, and C, respectively. As shown, the computing devices generate equivalent tags for the audio recording portions represented by a, b, d, e, and f. However, for the audio recording portion c, only computing devices B and C generate equivalent tags (see “dogs”), while computing device A generates non-equivalent tags (see “frogs”).

[0117] In some embodiments, only one transcriptional label (from the conflicting labels) is selected to be included in the training data, and majority voting can be applied to determine which label to include. For example, since two computing devices generated "dogs" for audio part c, while only one computing device generated "frogs", the majority winner is "dogs". In some embodiments, where weighting is applied among the labels, the "dogs" from computing device B and the "dogs" from computing device C are given equal weights, while the "frogs" from computing device A are given a smaller weight.

[0118] When non-equivalent labels are identified, a vote can be automatically conducted by the calculation system. Alternatively, the vote can be performed by a third-party human labeler.

[0119] In some cases, a particular computing device is known to produce more accurate transcription tags. Therefore, all transcription tags produced from that particular computing device can be more heavily weighted to be included in the underlying facts of the training data.

[0120] In other embodiments, all or more conflicting labels are included in the training data, wherein different accuracy probability weights are associated with different labels that will be consumed and processed by the trained model.

[0121] Figure 9 Additional or alternative methods for resolving conflicts involving non-equivalent transcription tags are illustrated. For example, utterance 910 is shown alongside its corresponding graphical representation with audio recording 920. A portion 930 of the audio recording is represented by the letter tags a, b, c, d, e, and f. The utterance is divided into one or more sets of microsegments; for example, utterance 910 is divided into three sets of microsegments, each set of which is distributed to multiple computing devices, such as computing device A 940, computing device B 950, and computing device C 960. All computing devices shown generate equivalent tags for audio recording portions a, b, d, e, and f (see tags 942, 952, 953). However, each computing device generates a non-equivalent tag for audio recording portion “c” (see “frogs” vs “dogs” vs “dogs”).

[0122] In some embodiments, where only a single transcription tag is selected to be included in the reconstructed electronic content, and / or when multiple transcription tags are included in the training and / or evaluation data, context weighting may be applied to determine the best tag (i.e., the tag most likely to be accurate for the corresponding audio portion and / or word of the utterance).

[0123] In this example, referring to computing device B, since the label for part c is generated using microsegments, where audio part c is preceded and followed by another word "two dogs ran," the computing system (and / or the computing device and / or the human entity) will determine that the label provided by device B for the labeled part "c" has more context, or a greater context weight for the labeled part "c," causing the system to select device B's label "dogs" for part "c." In other words, computing device B has the greatest access to the context used to label part "c," with the context preceding and following part "c" in the received microsegments [b, c, d] for labeling. In contrast, other devices (A or C) do not have as much context for labeling part c from a single received microsegment; each device only has the context preceding or following part "c" in a separate microsegment it receives with part "c" (e.g., device A receives segment [a, b, c] while device C receives segment [c, d]). This is even more important when the device does not receive a complete set of micro-segments (or, consecutive micro-segments) for a utterance, but only a limited set of micro-segments of the entire utterance, which further restricts access to the corresponding context of any particular term / part of the utterance to be tagged.

[0124] The reasoning logic confirms the hypothesis that broader context is associated with a higher probability of generating accurate labels. For example, in the case of computing device A, "frogs" is likely accurate because the context of the preceding word "two" indicates multiples. However, in the case of computing device C, "dog" is likely more accurate than "frogs" because "dog" is generally more associated with the verb "ran" than "frogs." Furthermore, as shown in the figure, it is advantageous to choose transcription labels generated by computing device B (and / or to be more heavily weighted than other transcription labels for audio part c) because the context from the preceding word "two" is generally associated with multiples (implying the plural form of the noun), and "ran" is generally more associated with "dogs" than "frogs." Therefore, including transcription labels (multiple) generated by computing device B facilitates the generation of effective training data because these labels most closely match the original utterance 910.

[0125] In some embodiments, each transcription tag may receive a different weighted score due to the identified context. For example, transcription tag 942 receives a first weighted score, and transcription tag 952 and transcription tag 962 receive a third weighted score, wherein the weighted score for transcription tag 952 is higher than the weighted scores for the other transcription tags because of a greater degree of context (i.e., preceding and subsequent words) for the non-equivalent tag of audio portion c. In some embodiments, the weighted scores for transcription tags 942 and 962 are equal (i.e., only one of the preceding or subsequent words). In some embodiments, transcription tags 942 and 962 are not identical, wherein determining the preceding context (e.g., word, phrase, etc.) or the subsequent context provides greater or less context for the non-equivalent tag.

[0126] Now let's turn our attention to... Figure 10 It shows a flowchart 1000, which includes components related to those described in the above references. Figure 1 The computing system 130 can implement various actions associated with exemplary methods, and / or various actions associated with protecting data access to machine learning training data during the labeling process of the corresponding data.

[0127] like Figure 10 As shown, flowchart 1000 and the corresponding method include the action of a computing system receiving electronic content comprising raw data (action 1010). The computing system then determines a security level associated with the electronic content (action 1020). Subsequently, the computing system selectively divides the electronic content into multiple micro-segments according to their duration, the duration of which is selected based on the determined security level (action 1030). Subsequently, the computing system selectively and restrictively distributes the multiple micro-segments to multiple destination computing devices (action 1040). It is worth noting that the action of distributing micro-segments (action 1040) in some cases also includes restricting or otherwise constraining the distribution of micro-segments, such that only a predetermined number of micro-segments from the raw data (or, a predefined set of data from the raw data) are distributed to any one of the destination computing devices.

[0128] After distribution, the computing system applies a corresponding tag (i.e., transcription tag) to each of the multiple destination computing devices (action 1060). The computing system then reconstructs the micro-segments (and their corresponding tags) into reconstructed electronic content, which includes training data for a machine learning model. Finally, the computing system trains the machine learning model with the reconstructed electronic content (action 1070). This training may include applying context and / or accuracy probability weights, which are included in or associated with different tags on the reconstructed micro-segments, particularly when multiple tags are provided for the same portion of the electronic content used in the training data.

[0129] It should also be understood that, in addition to using the reconstructed training data to train the model, in some cases, the reconstructed training data is also used to fine-tune the already trained machine learning model. In some embodiments, the machine learning model is trained for speech recognition, optical character recognition, and / or natural language processing.

[0130] Additionally, in some cases, it is anticipated that the reconstructed electronic content, including micro-segments and corresponding tags, will be further processed and / or modified before including or generating machine learning training data.

[0131] Furthermore, this method can be implemented by a computer system including one or more processors and computer-readable media (such as computer memory). Specifically, the computer memory can store computer-executable instructions that, when executed by one or more processors, cause various functions to be performed, such as the actions described in the embodiments.

[0132] Embodiments of the present invention may include or utilize dedicated or general-purpose computers including computer hardware, as discussed in more detail below. Embodiments within the scope of the present invention also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available medium accessible by a general-purpose or dedicated computer system. Computer-readable media storing computer-executable instructions are physical storage media. Computer-readable media carrying computer-executable instructions are transmission media. Therefore, by way of example and not limitation, embodiments of the present invention may include at least two distinct types of computer-readable media: physical computer-readable storage media and transmission computer-readable media.

[0133] Physical computer-readable storage media include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices (such as CD, DVD, etc.), magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store program code in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.

[0134] A “network” is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately considers that connection as a transmission medium. Transmission media may include networks and / or data links, which can be used to carry or transmit desired program code in the form of computer-executable instructions or data structures, and can be accessed by general-purpose or special-purpose computers. Combinations of the above are also included within the scope of computer-readable media.

[0135] Furthermore, upon arrival at various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission computer-readable medium to the physical computer-readable storage medium (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., a "NIC") and then ultimately transferred to the computer system RAM and / or a less lossy computer-readable physical storage medium at the computer system. Therefore, computer-readable physical storage media can be included in computer system components that also (or even primarily) utilize the transmission medium.

[0136] Computer-executable instructions include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a particular function or group of functions. Computer-executable instructions can be, for example, binary, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. Rather, the described features and actions are disclosed as exemplary forms for implementing the claims.

[0137] Those skilled in the art will understand that this invention can be practiced in network computing environments with various types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframes, mobile phones, PDAs, pagers, routers, switches, etc. This invention can also be implemented in distributed system environments, where both local and remote computer systems, linked via a network (or via a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links), perform tasks. In a distributed system environment, program modules can reside in both local and remote memory storage devices.

[0138] Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), program-specific integrated circuits (ASICs), program-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0139] The invention may be embodied in other specific forms without departing from its spirit or characteristics. The described embodiments are to be considered illustrative rather than restrictive in all respects. Therefore, the scope of the invention is defined by the appended claims rather than the foregoing description. All variations within the meaning and scope of equivalents of the claims are covered by the claims.

Claims

1. A method implemented by a computing system for protecting data access of machine learning training data at a plurality of distributed computing devices, the method comprising: receiving, by the computing system, electronic content comprising raw data; determining, by the computing system, a security level associated with the electronic content; selectively partitioning, by the computing system, the electronic content into a plurality of micro-segments according to a duration, the duration of the micro-segments selected according to the determined security level; identifying, by the computing system, a plurality of destination computing devices configured to apply a plurality of labels corresponding to the plurality of micro-segments; and distributing, by the computing system, the plurality of micro-segments restrictively to the plurality of destination computing devices, wherein the restrictive distribution comprises limiting a number of micro-segments to be distributed to any of the destination computing devices to less than a predetermined threshold, wherein the predetermined threshold is determined based on the determined security level. partitioning the electronic content into the plurality of micro-segments, wherein a first micro-segment has content that overlaps at least a portion of a second micro-segment.

2. The method of claim 1, wherein the dividing, by the computing system, the electronic content into a plurality of micro-segments further comprises:

3. The method of claim 2, wherein the first micro-segment has content that overlaps the second micro-segment by a predetermined number of words, the number of words identified by the computing system during the partitioning of the electronic content into at least the first micro-segment and the second micro-segment.

4. The method of any one of claims 1-3, further comprising: determining, by the computing system, that a micro-segment is to be further subdivided into two or more micro-segment fragments based on detecting a set of one or more attributes of the micro-segment, the set of one or more attributes of the micro-segment corresponding to an increased data security level compared to the determined security level.

5. The method of claim 4, wherein the detected set of one or more attributes of the micro-segment comprises one or more of: a sequence of numbers, one or more names, one or more characters or symbols identified as being non- normal to speech, or a specific term comprising an account number, a bank, a credit card, or other word or phrase determined to indicate data associated with a user or account credential. modifying the micro-segment by one or more of:

6. The method of any one of claims 1-3, further comprising: truncating the micro-segment at a natural speech break identified by the computing system; removing a local word identified in the micro-segment by the computing system; or changing a tone of speech audio included in the micro-segment.

7. The method of any one of claims 1-3, wherein the computing system restrictively distributing the plurality of micro-segments to the plurality of destination computing devices further comprises: accessing, by the computing system, a lookup table comprising identifiers, wherein each identifier corresponds to a destination computing device of the plurality of destination computing devices; determining, by the computing system, a number of micro-segments that have been distributed to each of the destination computing devices; and ​ the computing system links the identifier to the number of micro-segments that have been distributed to each of the destination computing devices, wherein the number of micro-segments corresponds to a percentage of the total micro-segments divided from the electronic content that is distributed to each of the destination computing devices.

8. The method of any of claims 1-3, wherein the computing system restrictively distributing the plurality of micro-segments to the plurality of destination computing devices further comprises: the computing system determining a number of micro-segments that have been distributed to a first destination computing device; the computing system identifying a particular micro-segment to be distributed to any destination computing device; the computing system determining that distributing the micro-segment to the first destination computing device causes a total number of micro-segments distributed to the first destination computing device to exceed the predetermined threshold; the computing system determining a number of micro-segments that have been distributed to a second destination computing device; the computing system determining that distributing the particular micro-segment to the second destination computing device does not cause the total number of micro-segments distributed to the second destination computing device to exceed the predetermined threshold; and the computing system distributing the particular micro-segment to the second destination computing device.

9. The method of any of claims 1-3, wherein the computing system restrictively distributing the plurality of micro-segments to the plurality of destination computing devices further comprises: the computing system determining a number of micro-segments that have been distributed to a first destination computing device; the computing system identifying a particular micro-segment to be distributed to one of the destination computing devices of the plurality of destination computing devices; the computing system determining that the particular micro-segment is contiguous with a predetermined number of micro-segments of the number of micro-segments that have been distributed to the first destination computing device; the computing system determining a number of micro-segments that have been distributed to a second destination computing device; the computing system determining that the particular micro-segment is non-contiguous with a predetermined number of micro-segments of the number of micro-segments that have been distributed to the second destination computing device; and the computing system distributing the particular micro-segment to the second destination computing device.

10. The method of any of claims 1-3, wherein the computing system restrictively distributing the plurality of micro-segments to the plurality of destination computing devices further comprises: the computing system identifying a particular micro-segment to be distributed to one of the destination computing devices of the plurality of destination computing devices; the computing system determining that distributing the particular micro-segment to a first destination computing device does not cause a total number of micro-segments distributed to the first destination computing device to exceed the predetermined threshold; the computing system distributing the particular micro-segment to the first destination computing device; the computing system determines that distributing the particular microsegment to a second destination computing device would not cause the total number of microsegments distributed to the second destination computing device to exceed the predetermined threshold; and the computing system distributes the particular microsegment to the second destination computing device.

11. The method of any of claims 1-3, further comprising: the computing system receiving a plurality of tags corresponding to a plurality of microsegments partitioned from the electronic content, the plurality of tags provided by at least two of the plurality of destination computing devices.

12. The method of claim 11, further comprising: the computing system reconstructing the plurality of microsegments into reconstructed electronic content comprising the plurality of tags corresponding to the plurality of microsegments, the reconstructed electronic content comprising training data for a machine learning model; and the computing system training the machine learning model with the reconstructed electronic content comprising training data for the machine learning model.

13. The method of claim 12, wherein reconstructing the plurality of microsegments into reconstructed electronic content further comprises microsegments partitioned into portions comprising microsegments that overlap one another, the method further comprising: the computing system determining that at least one overlapping portion of microsegments comprises a non-equivalent set of corresponding tags; and the computing system, for the overlapping portion of microsegments, including a particular tag of the non-equivalent set of corresponding tags in the reconstructed electronic content.

14. The method of claim 13, wherein the particular tag of the set of non-equivalent corresponding tags is selected for inclusion in the reconstructed electronic content based on a determination that the particular tag for the at least one overlapping portion is provided for a first microsegment having an increased level of context relative to a second microsegment having a different tag for the at least one overlapping portion.

15. A computing system configured for protecting data access of machine learning training data at a plurality of distributed computing devices, wherein the computing system comprises: one or more processors; and one or more computer-readable hardware storage devices storing computer-executable instructions structured to be executable by the one or more processors to cause the computing system to at least: receive electronic content comprising raw data corresponding to a determined level of security; selectively partition the electronic content into a plurality of microsegments by duration, the duration of the microsegments selected according to the determined level of security associated with the electronic content; identify a plurality of destination computing devices configured to apply a plurality of tags corresponding to the plurality of microsegments; distributing the plurality of micro-segments to the plurality of destination computing devices restrictively, wherein the restrictive distribution comprises limiting a number of micro-segments to be distributed to any of the plurality of destination computing devices to less than a predetermined threshold, wherein the predetermined threshold is determined based on the determined security level; receiving a plurality of tags corresponding to the plurality of micro-segments partitioned from the electronic content, the plurality of tags provided by at least two of the plurality of destination computing devices; and reconstructing the plurality of micro-segments into reconstructed electronic content comprising the plurality of tags corresponding to the plurality of micro-segments, the reconstructed electronic content comprising training data for a machine learning model.

Citation Information

Patent Citations

  • Multimodal transmission of packetized data

    CN108541312A

  • Automatically generating instructions from tutorials for search and user navigation

    CN110096576A