Information processing device, information processing system, and program

The information processing device segments and secures meeting data by voice and text segments, addressing mixed confidentiality levels with user-specific access control, enhancing security and accuracy.

JP7767767B2Active Publication Date: 2025-11-12FUJIFILM BUSINESS INNOVATION CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021134070
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-19
Publication Date
2025-11-12
Estimated Expiration
2041-08-19

AI Technical Summary

Technical Problem

Existing methods for securing meeting minutes data fail to effectively manage varying levels of confidentiality within a single meeting, as they often treat all data uniformly, neglecting the mixed nature of content types with different security needs.

Method used

An information processing device that segments meeting data into voice and text segments, assigns security levels based on content analysis, and controls access and output based on user authority, morpheme type, and predefined tags.

Benefits of technology

Enhances security by allowing granular control over access to meeting data, protecting sensitive information with higher confidentiality levels and enabling accurate segmentation and access restriction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767767000001
    Figure 0007767767000001
  • Figure 0007767767000002
    Figure 0007767767000002
  • Figure 0007767767000003
    Figure 0007767767000003
Patent Text Reader

Abstract

To provide an information processing device, an information processing system and a program which can more practically perform security protection of contents of minutes of a meeting etc., as compared to a case where the security protection is performed in a unit of data of the minutes.SOLUTION: In a management server being an information processing device, a control unit 11 comprises a text processing unit, a voice data division unit, a text data division unit, a level assigning unit, an encryption unit, a decryption unit and an output control unit. The text processing unit provides text data obtained by converting voice data into text to the text data division unit and the voice data division unit. The text data division unit and the voice data division unit divide the text data into a plurality of voice segments. The level assigning unit assigns a security level to each of the plurality of voice segments, according to contents of the text data and the voice data of each of the plurality of voice segments. The encryption unit, the decryption unit and the output control unit control each output of the plurality of voice segments, according to the security level.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing system, and a program. [Background technology]

[0002] When a meeting or the like is held in a company or other organization, the minutes are typically stored in the form of audio data or text data in a shared folder or the like. Such minutes data may contain confidential information, but they are often stored without any measures in place to prevent access to the minutes data, which poses a security risk. To address this issue, there are methods that encrypt the minutes data itself and allow access only to specific individuals, and methods such as Patent Document 1 that output audio data with segments containing personal information obscured. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special Publication No. 2009-501942 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in a meeting, for example, while discussing content that should be widely shared within the company, the discussion may switch to a discussion of personnel matters including personal information, or confidential matters related to management decisions. In this way, multiple types of content with varying levels of confidentiality may be discussed in no particular order within a single meeting. For this reason, it is not practical to restrict access to the data of meeting minutes.

[0005] An object of the present invention is to make it possible to protect the security of the contents of minutes of a meeting or the like more practically than when the protection is done on a data basis of minutes. [Means for solving the problem]

[0006] The invention described in claim 1 includes a processor, wherein the processor divides text data obtained by converting voice data into text and the voice data into a plurality of voice segments, assigns a security level to each of the plurality of voice segments according to the contents of the text data and voice data of each of the plurality of voice segments, and determines the security level. and the level of authority associated with the accessing user. Depending on For each of the plurality of audio sections displayed in a selectable list, processing is performed to restrict the user from viewing the text data and listening to the audio data, and the selected before Recording Vocal interval The text data and the audio data This is an information processing device characterized by controlling each output. The invention described in claim 2 is the information processing device described in claim 1, characterized in that the processor assigns the security level to each of the plurality of speech sections according to the type of each of the plurality of morphemes that constitute each of the plurality of speech sections as the content. The invention described in claim 3 is the information processing device described in claim 2, characterized in that the processor controls the output of at least one of each of the multiple speech segments and each of the multiple morphemes depending on the security level. The invention described in claim 4 is the information processing device described in claim 3, characterized in that the processor controls the output by outputting the location of the morpheme in the text data containing the morpheme and the audio data in a predetermined output format. The invention described in claim 5 is the information processing device described in claim 2, characterized in that if the type of the morpheme is a proper noun or a number indicating predetermined content, the processor assigns a higher security level to the audio section containing the morpheme than to other audio sections. The invention described in claim 6 is the information processing device described in claim 1, characterized in that the processor assigns a security level that is preset for each type of content to each of the multiple voice segments. The invention described in claim 7 is the information processing device described in claim 6, characterized in that the processor pre-sets the security level in tag information indicating the type of content. The invention described in claim 8 is the information processing device described in claim 1, characterized in that the processor divides the text data and the audio data into the multiple audio segments using at least one of machine learning and natural language processing. The invention described in claim 9 is the information processing device described in claim 8, characterized in that the processor further divides the voice into the multiple voice segments by estimating the transition of the subject uttering the voice based on information on changes in voice characteristics obtained from the voice data. In the invention described in claim 10, the processor controls the output by associating a decryption key for decrypting the encrypted audio section to make it viewable with each of the security levels, Record Depending on the level of limitation, The aforementioned 2. The information processing device according to claim 1, wherein the decryption key is given to a user. The invention described in claim 11 is the information processing device described in claim 10, characterized in that the processor further assigns the decryption key to the user according to a predetermined range of the user. The invention described in claim 12 is the information processing device described in claim 10, characterized in that the processor further assigns the decryption key to the user based on at least one of the organization to which the user belongs and the role assigned to the user. The invention described in claim 13 is the information processing device described in claim 10, characterized in that the processor further assigns the decryption key to the user according to a predetermined period of time. The invention described in claim 14 is the information processing device described in claim 10, characterized in that when the user to whom the decryption key is assigned accesses the information processing device, the processor controls the output of at least one of text data and audio data of the audio section corresponding to the decryption key. The invention described in claim 15 includes a division means for dividing text data obtained by converting voice data into text and the voice data into a plurality of voice segments, an assignment means for assigning a security level to each of the plurality of voice segments in accordance with the contents of the text data and the voice data of each of the plurality of voice segments, and a security level assigning means for assigning a security level to each of the plurality of voice segments in accordance with the contents of the text data and the voice data of each of the plurality of voice segments. and the level of authority associated with the accessing user. Depending on For each of the plurality of audio sections displayed in a selectable list, processing is performed to restrict the user from viewing the text data and listening to the audio data, and the selected before Recording Vocal interval The text data and the audio data and an output control means for controlling each output. The invention described in claim 16 is a computer-implemented method for generating text data obtained by converting voice data into text and dividing the voice data into a plurality of voice segments, and a function of assigning a security level to each of the plurality of voice segments in accordance with the contents of the text data and the voice data of each of the plurality of voice segments, and a function of assigning a security level to each of the plurality of voice segments in accordance with the contents of the text data and the voice data of each of the plurality of voice segments. and the level of authority associated with the accessing user. Depending on For each of the plurality of audio sections displayed in a selectable list, processing is performed to restrict the user from viewing the text data and listening to the audio data, and the selected before Recording Vocal interval The text data and the audio data It is a program that realizes the function to control each output. [Effects of the Invention]

[0007] According to the present invention of claim 1, by protecting the security of data of minutes of meetings, etc. on a voice section basis, it is possible to provide an information processing device that enables more practical operation than when protecting the security of data on a minutes basis. According to the present invention of claim 2, it is possible to control the output of a speech segment depending on the type of morpheme that constitutes the speech segment. According to the present invention of claim 3, it is possible to control output not only for each speech section but also for each morpheme that constitutes the speech section. According to the present invention of claim 4, a user who listens to the audio data in the audio section can know the parts where security is protected. According to the present invention of claim 5, when a voice section contains highly confidential content, it is possible to protect security at a higher level than other voice sections. According to the present invention of claim 6, a security level can be automatically assigned simply by determining the type of content. According to the present invention of claim 7, it is possible to automatically assign a security level according to the content of the assigned tag information. According to the present invention of claim 8, for example, by dividing the voice section using machine learning or natural language processing by AI (artificial intelligence), it is possible to achieve division with higher accuracy than when dividing simply by the number of characters or time. According to the present invention of claim 9, segmentation of voice segments is performed taking into consideration the subject of the voice, thereby enabling segmentation with higher accuracy. According to the present invention of claim 10, the audio sections that a user can listen to can be restricted according to the user's authority. According to the present invention of claim 11, it is possible to limit the audio sections that users can listen to, not only according to the user's authority but also according to a predetermined range of users. According to the present invention of claim 12, the audio sections that a user can listen to can be restricted not only according to the user's authority but also according to the organization to which the user belongs and the role that the user has been assigned. According to the present invention of claim 13, the audio section that a user can listen to can be restricted not only according to the user's authority but also according to a predetermined period. According to the present invention of claim 14, authorized users can view and read the text data and audio data of the target audio section. As a result, it becomes possible to understand the content at a glance from the text data, and to grasp the nuances of the utterances and supplement the accuracy of the text data from the audio data. According to the present invention of claim 15, by protecting the security of data in minutes of meetings, etc., in units of voice sections, it is possible to provide an information processing system that enables more practical operation than when protecting the security of data in units of minutes. According to the present invention of claim 16, by protecting the security of data in minutes of meetings, etc., in units of voice sections, it is possible to provide a program that enables more practical operation than when protecting the security of data in units of minutes of meetings, etc. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram showing the overall configuration of an information processing system to which an embodiment of the present invention is applied; [Figure 2] FIG. 2 illustrates a hardware configuration of a management server. [Figure 3] FIG. 2 is a diagram illustrating a hardware configuration of a client terminal. [Figure 4] FIG. 2 is a diagram illustrating a functional configuration of a control unit of the management server. [Figure 5] 10 is a flowchart showing the flow of encryption processing by the management server. [Figure 6] 10 is a flowchart showing the flow of processing by the management server when a user views a content. [Figure 7] FIG. 10 is a conceptual diagram visualizing audio data that is the target of encryption processing by the management server. [Figure 8] FIG. 10 is a diagram showing a specific example in which a security level is assigned to each voice segment. [Figure 9] 1A is a diagram showing a specific example of a viewing authority level granted to each user, and FIG. 1B is a diagram showing a specific example of a viewing authority level granted to each user and a security level granted to each audio segment. [Figure 10] FIG. 10 is a diagram showing a specific example of a login screen among the screens displayed on the display unit of the client terminal. [Figure 11] FIG. 10 is a diagram showing a specific example of a conference selection screen among the screens displayed on the display unit of the client terminal. [Figure 12] FIG. 10 is a diagram showing a specific example of a conference details screen among the screens displayed on the display unit of a client terminal of a user whose viewing authority level is "high." [Figure 13] FIG. 10 is a diagram showing a specific example of a conference details screen among the screens displayed on the display unit of a client terminal of a user whose viewing authority level is "medium." [Figure 14] FIG. 10 is a diagram showing another specific example of the conference details screen among the screens displayed on the display unit of the client terminal of a user whose viewing authority level is "medium." [Figure 15] FIG. 10 is a diagram showing a specific example in which a disclosure range is further assigned to each speech segment. [Figure 16] FIG. 10 is a diagram showing a specific example of an organization to which a user belongs and a role assigned to the user, which are associated with each user. [Figure 17] FIG. 10 is a diagram showing a specific example in which a non-disclosure period is further added to each voice section. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. <Configuration of information processing system> FIG. 1 is a diagram showing the overall configuration of an information processing system 1 to which this embodiment is applied. The information processing system 1 is configured by connecting a management server 10 and client terminals 30-1 to 30-n (n is an integer value of 2 or more) via a network 90. ​​The network may be, for example, a LAN (Local Area Network), the Internet, or the like.

[0010] The management server 10 is an information processing device that serves as a server that manages the entire information processing system 1. For example, the management server 10 divides text data, which is speech data converted into text, into multiple speech segments and assigns a security level to each speech segment depending on the content of the text data and speech data for each speech segment. The "security level" refers to an index that indicates the degree of confidentiality of information. In this embodiment, there are three security levels, with "high," "medium," and "low" being assigned in descending order of security level. The management server 10 also controls the output of text data and speech data for speech segments by encrypting and decrypting the text data and speech data for speech segments.

[0011] Each of the client terminals 30-1 to 30-n is an information processing device such as a smartphone, personal computer, or tablet terminal used by each of users M1 to Mn who make up an organization G such as a company. Each of the client terminals 30-1 to 30-n, for example, records the audio of a conference and transmits the audio data to the management server 10. Also, for example, each of the client terminals 30-1 to 30-n accesses the management server 10 and outputs text data and audio data of the speech section. Hereinafter, when it is not necessary to distinguish between the client terminals 30-1 to 30-n and users M1 to Mn in the description, these will be collectively referred to as the client terminal 30 and user M.

[0012] Note that there are cases where each of the client terminals 30-1 to 30-n is used individually by each of the users M1 to Mn, and cases where multiple users M share one client terminal 30. For example, there are cases where an external microphone is connected to the client terminal 30-1 of the user M1, and the microphone is installed in a conference room to record a conference of the users M1 to M5.

[0013] <Hardware configuration of the management server> FIG. 2 is a diagram showing the hardware configuration of the management server 10. As shown in FIG. The management server 10 has a control unit 11, a memory 12, a storage unit 13, a communication unit 14, an operation unit 15, and a display unit 16. These units are connected by a data bus, an address bus, a PCI (Peripheral Component Interconnect) bus, etc.

[0014] The control unit 11 is a processor that controls the operation of the management server 10 through the execution of various software such as an OS (operating system) and application software. The control unit 11 is configured, for example, by a CPU (Central Processing Unit). The memory 12 is a storage area that stores various software and data used for executing the software, and is used as a working area for calculations. The memory 12 is configured, for example, by a RAM (Random Access Memory).

[0015] The storage unit 13 is a storage area that stores input data for various software programs and output data from various software programs, and stores a database that stores various information. The storage unit 13 is configured with, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), a semiconductor memory, etc. that are used to store programs, various setting data, etc. The communication unit 14 transmits and receives data via the network 90. ​​The communication unit 14 transmits and receives data to and from the client terminal 30.

[0016] The operation unit 15 is configured with, for example, a keyboard, a mouse, mechanical buttons, and switches, and accepts input operations. The operation unit 15 also includes a touch sensor that forms a touch panel integrally with the display unit 16. The display unit 16 displays images, text information, and the like. The display unit 16 is configured with, for example, a liquid crystal display or an organic EL (Electro Luminescence) display used to display information. <Client terminal hardware configuration> FIG. 3 is a diagram showing the hardware configuration of the client terminal 30. As shown in FIG. The client terminal 30 has a hardware configuration similar to that of the management server 10 shown in Fig. 2, except for a recording unit 37 that records audio. That is, the client terminal 30 has a control unit 31 configured with a processor such as a CPU, a memory 32 configured with a storage area such as RAM, and a storage unit 33 configured with a storage area such as an HDD, SSD, or semiconductor memory. The client terminal 30 also has a communication unit 34 that transmits and receives data to and from the management server 10 via the network 90, and an operation unit 35 configured with a keyboard, mouse, touch panel, etc. The client terminal 30 also has a display unit 36 ​​configured with a liquid crystal display, organic EL display, etc. These units are connected by a data bus, an address bus, a PCI bus, etc.

[0017] <Functional configuration of the control unit of the management server> FIG. 4 is a diagram showing the functional configuration of the control unit 11 of the management server 10. As shown in FIG. The control unit 11 of the management server 10 functions as an audio data acquisition unit 101, a text conversion unit 102, a text data classification unit 103, an audio data classification unit 104, a level assignment unit 105, an encryption unit 106, an access reception unit 107, a decryption unit 108, and an output control unit 109.

[0018] The voice data acquisition unit 101 acquires voice data transmitted from the client terminal 30. The voice data acquired by the voice data acquisition unit 101 is associated with information relating to time. This makes it possible to identify the elapsed time from the start of recording in morpheme units. The voice data acquired by the voice data acquisition unit 101 is stored in a database in the storage unit 13.

[0019] The text conversion unit 102 converts the voice data acquired by the voice data acquisition unit 101 into text. Specifically, the text conversion unit 102 generates text data by performing voice recognition on the voice data acquired by the voice data acquisition unit 101. In the voice recognition, for example, a method is used in which feature quantities such as the strength, frequency, and interval of sounds contained in the voice data are extracted, and the matching rate with phoneme and word models stored in advance in a database is calculated to recognize the data as words. The text data generated by the text conversion unit 102 is associated with the voice data before being converted into text and stored in the database of the storage unit 13.

[0020] The text data segmentation unit 103 segments the text data generated by the text generation unit 102 into multiple contexts. Each segmented context is called a speech segment. Methods for segmenting text data into multiple speech segments include, for example, a method of segmenting by setting a threshold for the number of characters or time of the text data, a machine learning method using AI (artificial intelligence), and a natural language processing method.

[0021] The audio data segmentation unit 104 segments the audio data before it is converted into text by the text conversion unit 102 into a plurality of audio segments according to the result of segmentation by the text data segmentation unit 103. Specifically, for example, the audio data segmentation unit 104 segments the audio data into a plurality of audio segments by cutting out the audio data based on the start time and end time of each of the plurality of audio segments of the text data segmented into the plurality of audio segments. This results in the text data and the audio data corresponding to the same audio segment.

[0022] Furthermore, the voice data segmentation unit 104 can directly segment the voice data into a plurality of voice segments, regardless of the result of segmenting the text data, or as a complement to the result of segmenting the text data. For example, it can estimate the transition of the subject uttering the voice from changes in voice characteristics obtained from the voice data, and segment the voice data into a plurality of voice segments based on the estimation result.

[0023] For example, from the voice data of a conference or the like in which users M1 to M5 participated, differences in the voice characteristics of each of users M1 to M5 can be obtained, and therefore it is possible to estimate which of users M1 to M5 is the subject of the voice. Note that when the voice data is directly divided into multiple voice segments without relying on the results of dividing the text data, adjustments are made with the results of dividing the text data. As a result, the text data and the voice data correspond to the same voice segment.

[0024] The level assigning unit 105 assigns a security level to each of the multiple voice segments according to the content of the text data and voice data of each of the multiple voice segments. In this embodiment, the security levels are assigned in descending order of "high," "medium," and "low." The level assigning unit 105 can also assign a security level in advance to each piece of tag information that is assigned according to the type of content of the text data and voice data of the voice segment.

[0025] The content type of the text data and speech data of the speech section is identified by the type of each of the multiple morphemes that make up the text data and speech data of the speech section. Morphemes are, for example, proper nouns or numbers. This allows a security level to be automatically assigned simply by identifying tag information for the speech section. Specific examples of tag information will be described later with reference to Figures 7 and 8, etc.

[0026] The encryption unit 106 encrypts the text data and audio data of each of the multiple audio segments according to the security level assigned to each audio segment by the level assigning unit 105. A decryption key that decrypts the text data and audio data of each of the multiple audio segments encrypted by the encryption unit 106 and makes the data viewable is associated with each security level. This makes it possible to control the output for each audio segment that makes up the audio data. A specific example of encryption for each of the multiple audio segments will be described later with reference to FIG. 7.

[0027] The encryption unit 106 can also encrypt each of the multiple morphemes that make up the speech section. In this case, a decryption key that decrypts the text data and audio data containing the multiple morphemes encrypted by the encryption unit 106 and makes them viewable is associated with each security level. This makes it possible to control the output for each morpheme that makes up the speech section. A specific example of encrypting each of the multiple morphemes that make up the speech section will be described later with reference to FIG. 14.

[0028] The access receiving unit 107 receives access from the client terminal 30 to the text data and audio data stored in the database of the storage unit 13 . The decryption unit 108 grants a decryption key to the user M of the accessing client terminal 30 in accordance with the viewing authority level previously granted to the user M and the security level granted to each audio section, thereby enabling the user M to view the text data and audio data of the corresponding audio section.

[0029] In this embodiment, viewing authority levels of "high," "medium," and "low" are assigned. A user M assigned a viewing authority level of "high" is assigned a decryption key that enables viewing of all audio sections in which text data and audio data are encrypted. A user M assigned a viewing authority level of "medium" is assigned a decryption key that enables viewing of audio sections in which text data and audio data are encrypted and that have security levels of "medium" and "low." A user M assigned a viewing authority level of "low" is assigned a decryption key that enables viewing of audio sections in which text data and audio data are encrypted and that have security level of "low."

[0030] Furthermore, the decryption unit 108 assigns a decryption key to the user M according to a predetermined range of the user M. Specifically, the decryption unit 108 assigns a decryption key to the user M according to, for example, at least one of the organization to which the user M belongs and the role assigned to the user M. This enables the user M to listen to the text data and audio data of the corresponding voice section. Note that a specific example of assigning a decryption key to the user M according to a predetermined range of the user M will be described later with reference to FIGS. 15 and 16.

[0031] Furthermore, the decryption unit 108 assigns a decryption key to the user M according to a predetermined period. Specifically, the decryption unit 108 can assign a decryption key to all users M, for example, after a non-disclosure period predetermined for each piece of tag information has elapsed. Note that a specific example of the non-disclosure period predetermined for each piece of tag information will be described later with reference to FIG. 17.

[0032] The output control unit 109 controls the display unit 16 of the client terminal 30 to display the text data and speech data encrypted or decrypted for each speech section or for each morpheme. Specifically, when a user M to whom a decryption key has been assigned accesses the client terminal 30, the output control unit 109 controls the output of at least one of the text data and speech data for the speech section corresponding to the decryption key. Specific examples of the text data and speech data displayed on the display unit 16 of the client terminal 30 will be described later with reference to FIGS. 12 to 14.

[0033] <Administration Server Processing> 5 and 6 are flowcharts showing the flow of processing by the management server 10. FIG. FIG. 5 shows the flow of encryption processing by the management server 10. When voice data is transmitted from the client terminal 30 (YES in step 401), the management server 10 acquires the voice data (step 402) and converts it into text (step 403). Specifically, the management server 10 generates text data by performing voice recognition on the acquired voice data. On the other hand, if voice data is not transmitted from the client terminal 30 (NO in step 401), the management server 10 repeats the processing of step 401 until voice data is transmitted from the client terminal 30.

[0034] The management server 10 divides the generated text data into a plurality of voice segments (step 404), and divides the voice data into a plurality of voice segments according to the division results (step 405). Thereafter, the management server 10 assigns a security level to each voice segment (step 406), and encrypts each voice segment according to the security level (step 407). This completes the encryption process by the management server 10.

[0035] FIG. 6 shows the flow of processing by the management server 10 when user M views a content. If the management server 10 receives access from the client terminal 30 (YES in step 501) and the viewing authority level of the accessing user M is "high" (YES in step 502), the management server 10 decrypts all encrypted voice segments (step 503). The management server 10 then controls to output the decrypted voice segments (step 507). This makes it possible to view the text data and voice data of all voice segments. On the other hand, if there is no access from the client terminal 30 (NO in step 501), the management server 10 repeats the processing of step 501 until there is access from the client terminal 30.

[0036] When the management server 10 receives access from the client terminal 30 (YES in step 501) and the viewing authority level of the accessing user M is not "high" (NO in step 502) but "medium" (YES in step 504), the management server 10 decrypts the voice segments with security levels of "medium" and "low" (step 505). The management server 10 then controls to output the decrypted voice segments (step 507). This makes it possible to view the text data and audio data of the voice segments with security levels of "medium" or "low."

[0037] On the other hand, if the viewing authority level of the accessing user M is neither "high" nor "medium" (NO in steps 502 and 504), the management server 10 determines that the accessing user's viewing authority level is "low" and decrypts the voice section with a security level of "low" (step 506). Then, the management server 10 controls to output the decrypted voice section (step 507). As a result, the text data and audio data of the voice section with a security level of "low" become viewable, and the processing of the management server 10 ends.

[0038] <Example> FIG. 7 is a conceptual diagram that visualizes the audio data that is the target of encryption processing by the management server 10. In FIG. As described above, the voice data acquired by the management server 10 is converted into text and divided into a plurality of voice segments. FIG. 7 shows an example in which the text data obtained by converting the voice data into text is divided into voice segments 1 to N (N is an integer value of 3 or more). Also shown is an example in which the voice data is divided into voice segments 1 to N in accordance with the division of the text data. Each of the voice segments 1 to N is assigned a tag 1 to N as tag information, and each of the tags 1 to N is assigned a security level. For example, each of the tags 1, 2, and N is assigned a security level of "low," "medium," and "high," respectively.

[0039] As described above, the management server 10 performs encryption according to the security level assigned to each voice section, and associates a decryption key for decrypting the encrypted text data and voice data with each voice section. In the example of Fig. 7, decryption keys 3, 2, and 1 are associated with each of the encrypted text data of voice sections 1 to N. Also, decryption keys C, B, and A are associated with each of the encrypted voice data of voice sections 1 to N.

[0040] FIG. 8 is a diagram showing a specific example in which a security level is assigned to each voice section. As shown in FIG. 8, if the content of the text data and audio data of the audio section is, for example, the date and time, venue, participants, and agenda of a meeting, tag information called "meeting information" is assigned. A security level of "low" is assigned to the tag information. This tag information applies to all meetings. If the content of the text data and audio data of the audio section is, for example, sales, operating profit, mergers and acquisitions (M&A) information, tag information called "management information" is assigned. A security level of "high" is assigned to the tag information. This tag information applies to, for example, management meetings.

[0041] Furthermore, if the content of the text data and voice data of the voice section is, for example, a person being evaluated, a performance review, or a new post, tag information called "personnel information" is assigned. A security level of "high" is assigned to this tag information. Such tag information is applicable, for example, to management meetings and department management meetings. Furthermore, if the content of the text data and voice data of the voice section is, for example, a product name, development code, market introduction date, cost, etc., tag information called "product information" is assigned. A security level of "high" is assigned to this tag information. Such tag information is applicable, for example, to product planning meetings, development proposal meetings, design reviews, etc.

[0042] Furthermore, if the content of the text data and voice data in the voice section is, for example, a sales destination, a business negotiation scale, a delivery schedule, etc., tag information called "sales information" is assigned. A security level of "high" is assigned to this tag information. Such tag information is targeted, for example, at a sales meeting. Furthermore, if the content of the text data and voice data in the voice section is, for example, a person's name, date of birth, contact information, address, etc., tag information called "personal information" is assigned. A security level of "high" is assigned to this tag information. Such tag information is targeted, for example, at a human resources meeting.

[0043] Furthermore, if the content of the text data and voice data of the voice section is, for example, an assignment, tag information called "Assignment" is assigned. A security level of "Medium" is assigned to the tag information. This type of tag information applies to all meetings. Furthermore, if the content of the text data and voice data of the voice section is, for example, an action item, a person in charge, or a deadline, tag information called "ToDo" is assigned. A security level of "Medium" is assigned to the tag information. This type of tag information applies to all meetings.

[0044] FIG. 9A is a diagram showing a specific example of the viewing authority level given to each user M. 9(A) illustrates users M1 to M3 as accessors to the text data and audio data stored in the management server 10. Of these, user M1 is given a viewing authority level of "high" in advance. In this case, user M1 is given decryption keys 1 to 3 for the text data and decryption keys A to C for the audio data shown in FIG. 7.

[0045] Also, user M2 is given a viewing authority level of "medium" in advance. In this case, user M2 is given decryption keys 2 and 3 for the encrypted text data and decryption keys B and C for the encrypted audio data in Fig. 7. Also, user M3 is given a viewing authority level of "low" in advance. In this case, user M3 is given decryption key 3 for the encrypted text data and decryption key C for the encrypted audio data in Fig. 7.

[0046] FIG. 9B is a diagram showing a specific example of the viewing authority level given to user M and the security level given to each voice segment. 9(B) illustrates users M1 to M3 as accessors to the text data and audio data stored in the management server 10. Of these, user M1 has been granted a viewing authority level of "high" in advance. In this case, user M1 is able to view the text data and audio data of all audio segments.

[0047] Also, user M2 is given a viewing authority level of "medium" in advance. In this case, user M2 can view text data and audio data of the audio section with a security level of "low" or "medium." Also, user M3 is given a viewing authority level of "low." In this case, user M3 can view text data and audio data of the audio section with a security level of "low."

[0048] 10 to 14 are diagrams showing specific examples of screens displayed on the display unit 36 ​​of the client terminal 30. FIG. FIG. 10 shows a specific example of a login screen among the screens displayed on the display unit 36 ​​of the client terminal 30. When user M operates the client terminal 30 to access text data and audio data of a conference or the like stored in the database of the storage unit 13 of the management server 10, the login screen shown on the left side of FIG. 10 is first displayed on the client terminal 30. User M enters "ID" and "password" in the input fields provided on the login screen and presses the button B1 labeled "Login." This displays the screen shown on the right side of FIG. 10. This screen displays the viewing authority level previously granted to user M who performed the login operation. In the example of FIG. 10, the viewing authority level previously granted to user M is "High." This allows user M who performed the login operation to immediately know his / her own viewing authority level. After knowing his / her own viewing authority level, user M presses the button B2 labeled "Next."

[0049] FIG. 11 shows a specific example of a conference selection screen among the screens displayed on the display unit 36 ​​of the client terminal 30. When user M presses button B2 labeled "Next" in FIG. 10, a list of conferences in which user M can listen to text data and audio data is displayed in a selectable manner. FIG. 11 displays information about each item of the conference name, date, location, and main participants as a list of conferences. For example, a conference named "Product X Design Review" is shown to have been held on "2021 / 1 / 12," at "Conference Room B, Office A," and with "Tanaka, Sato, and 10 others" as its main participants.

[0050] Additionally, the meeting titled "Product Y Development Start Proposal" was held on "1 / 18 / 2021," at "Head Office Conference Room C," and with main participants "Suzuki, Mori, and 8 others." Additionally, the meeting titled "Product Z Market Trouble Countermeasures Meeting" was held on "1 / 18 / 2021," at "Business Office A Conference Room D," and with main participants "Yamada, Takahashi, and 7 others."

[0051] In the example of FIG. 11, three conferences that user M can view are shown as examples, but this is merely an example. For example, by pressing button B3 labeled "Previous Page" or button B4 labeled "Next Page," more conferences that user M can view are displayed. Also, as shown in FIG. 11, user M can search for conferences by keyword. Here, assume that user M, whose viewing authority level is "High," presses button B5 labeled "View" located on the right side to select the conference whose name is "Product X Design Review." Then, the conference details screen shown in FIG. 12 is displayed on the client terminal 30.

[0052] FIG. 12 shows a specific example of a conference details screen among the screens displayed on the display unit 36 ​​of the client terminal 30 of user M, whose viewing authority level is "high." When user M, whose viewing authority level is "high," presses the button B5 labeled "Watch" in FIG. 11 to select the conference titled "Product X Design Review," a conference details screen shown in FIG. 12 is displayed. This conference details screen displays the security level for each audio section of the conference titled "Product X Design Review," the elapsed time since the start of the conference, the contents of the text data, and a button B8 labeled "Play" for listening to the audio data. The example in FIG. 12 is a screen displayed on the client terminal 30 of user M, whose viewing authority level is "high," so all audio sections are displayed in a viewable manner. User M reads the text data and listens to the audio data as necessary. This makes it easier to understand, for example, the emotions of participants and the nuances of their words, which cannot be grasped from text data.

[0053] The example in Figure 12 shows an audio section where the security level is "low," the time elapsed since the start of the meeting is "0:00:00 to 0:01:23," and the content of the text data is "Today is January 12, 2021. The location is Conference Room B at Business Office A. The participants are Tanaka, Sato, and..."; an audio section where the security level is "high," the time elapsed since the start of the meeting is "0:01:23 to 0:02:34," and the content of the text data is "This component C requires high strength, so material D has been selected and the target cost is E yen,"; and an audio section where the security level is "medium," the time elapsed since the start of the meeting is "0:02:34 to 0:03:45," and the content of the text data is "Mr. F, please conduct further research on G and submit a report by next Wednesday."

[0054] In the example of FIG. 12, three audio segments that user M can view are shown as examples, but this is merely an example, and by pressing, for example, button B6 labeled "previous page" or button B7 labeled "next page," more audio segments that user M can view are displayed. Also, as shown in FIG. 12, user M can search for audio segments by keyword. As shown in FIG. 12, user M, whose viewing authority level is "high," can view text data and audio data of all audio segments without any restrictions.

[0055] FIG. 13 shows a specific example of a conference details screen among the screens displayed on the display unit 36 ​​of the client terminal 30 of user M, whose viewing authority level is "medium." When user M, whose viewing authority level is "medium," presses button B5 labeled "Watch" in Fig. 11 to select the conference named "Product X Design Review," the conference details screen shown in Fig. 13 is displayed. This conference details screen displays the security level for each audio section of the conference named "Product X Design Review," the elapsed time since the start of the conference, the contents of the text data, and button B8 labeled "Play" for listening to the audio data.

[0056] 13 is a screen displayed on the client terminal 30 of user M, whose viewing authority level is "medium," and therefore audio segments with a security level of either "low" or "medium" are displayed in a viewable manner. Although audio segments with a security level of "high" are displayed on the conference details screen, the text data is masked by blurring or blacking out to prevent it from being seen, and normal audio is not played back even when the button B8 labeled "play" is pressed. Furthermore, although not shown, a button labeled "play not allowed" may be displayed for audio data of audio segments with a security level of "high," preventing user M from even pressing the button.

[0057] FIG. 14 shows another specific example of the conference details screen among the screens displayed on the display unit 36 ​​of the client terminal 30 of user M, whose viewing authority level is "medium." When user M, whose viewing authority level is "medium," presses button B5 labeled "view" in Fig. 11 to select a conference named "Product X Design Review," a conference details screen shown in Fig. 14 is displayed. This conference details screen displays the security level for each audio section of the conference named "Product X Design Review," the elapsed time since the start of the conference, the contents of the text data, and a button B8 labeled "play" for listening to the audio data.

[0058] 14 is a screen displayed on the client terminal 30 of user M, whose viewing authority level is "medium," and therefore audio segments with a security level of either "low" or "medium" are displayed in a viewable manner. Note that, unlike in FIG. 13, audio segments with a security level of "high" are masked so that part of the text data cannot be seen.

[0059] In this way, instead of treating all text data of a single speech section in the same way, masking can be performed on a morpheme-by-morpheme basis. This is because, as described above, a security level can be assigned to the speech section depending on the type of each of the multiple morphemes that make up the speech section. In this case, the masked portions are those portions of the text data of the speech section that have been determined to be important elements. Examples of important elements include proper nouns related to privacy, such as people's names and company names, and numerical values ​​related to trade secrets, such as addresses, dates, and prices. In addition, the speech data is processed so that the portions determined to be important elements cannot be heard. For example, silence or a noise effect such as a beep is applied to the portions determined to be important elements.

[0060] FIG. 15 is a diagram showing a specific example in which a disclosure range is further assigned to each speech section. 8 above, a specific example in which a security level is assigned to tag information for each audio section has been described, but tag information can also be associated with a disclosure range. The "disclosure range" refers to a restriction that can be set in addition to the viewing authority level for those who can view the text data and audio data of the audio section.

[0061] In the example of FIG. 15, when the tag information is "conference information," the disclosure range is set to "ALL." This indicates that there are no restrictions on who can view the text data and audio data of the audio section other than the viewing authority level. Also, when the tag information is "management information," the disclosure range is set to "executives." This indicates that there is a restriction on who can view the text data and audio data of the audio section, other than the viewing authority level, that is, only user M, who is an executive.

[0062] Furthermore, when the tag information is "personnel information," the disclosure scope is set to "all users in the personnel department / members of the department management committee." This indicates that, in addition to the viewing authority level, there is a restriction that the people who can view the text data and audio data of the audio section are limited to all users M who belong to the personnel department and users M who are members of the department management committee. Furthermore, when the tag information is "product information," the disclosure scope is set to "all users in the design department." This indicates that, in addition to the viewing authority level, there is a restriction that the people who can view the text data and audio data of the audio section are limited to all users M who belong to the design department.

[0063] Furthermore, when the tag information is "Sales Information," the disclosure scope is set to "Sales Department ALL." This indicates that, in addition to the viewing authority level, there is a restriction on who can view the text data and audio data of the audio section, limiting it to all users M who belong to the sales department. Furthermore, when the tag information is "Personal Information," the disclosure scope is set to "Human Resources Department ALL / Executives." This indicates that, in addition to the viewing authority level, there is a restriction on who can view the text data and audio data of the audio section, limiting it to users M who belong to the human resources department and users M who are executives. Furthermore, when the tag information is "Task," and when the tag information is "ToDo," the disclosure scope is set to "ALL." This indicates that, in addition to the viewing authority level, there is no restriction on who can view the text data and audio data of the audio section.

[0064] FIG. 16 is a diagram showing a specific example of the organization to which the user M belongs and the role assigned to the user M, which are associated with each user M. In FIG. 15 shows the disclosure scope, which is a restriction imposed in addition to the viewing authority level. Therefore, in the management server 10, information about user M is managed in association with information about the organization to which the user belongs and the role to which the user is assigned, in addition to the viewing authority level. Note that this information may be acquired, for example, from a core system or the like managed separately within the company.

[0065] 16 shows the viewing authority level, department, and job title for each of users M1 to M4 who access the text data and audio data stored in the management server 10. For example, user M1's viewing authority level is "high," his / her organizational department is "design department," and his / her role is "manager." User M2's viewing authority level is "medium," his / her organizational department is "personnel department," and his / her role is "general" (general employee). User M3's viewing authority level is "low," his / her organizational department is "sales department," and his / her role is "general." User M4's viewing authority level is "high," his / her organizational department is "research laboratory," and his / her role is "director."

[0066] FIG. 17 is a diagram showing a specific example in which a non-disclosure period is further added to each voice section. 15, it has been explained that the disclosure range can be further associated with the tag information for each audio section, but a non-disclosure period can also be further associated. The "non-disclosure period" refers to a restriction that can be set on the period during which the text data and audio data of the audio section cannot be viewed, and once the non-disclosure period has elapsed, all users M will be able to view the data.

[0067] In the example of FIG. 17, when the tag information is "conference information," the non-disclosure period is set to "none." This indicates that there is no limit on the period during which the text data and audio data of the audio section cannot be viewed. Furthermore, when the tag information is "management information," the non-disclosure period is set to "two years." This indicates that after the two-year non-disclosure period has passed, all users M will be able to view the data.

[0068] Furthermore, when the tag information is "personnel information," the non-disclosure period is set to "3 years." This indicates that all users M will be able to view the content after the non-disclosure period of 3 years has passed. Furthermore, when the tag information is "product information," the non-disclosure period is set to "2 years." This indicates that all users M will be able to view the content after the non-disclosure period of 2 years has passed. Furthermore, when the tag information is "sales information," the non-disclosure period is set to "1 year." This indicates that all users M will be able to view the content after the non-disclosure period of 1 year has passed.

[0069] Furthermore, when the tag information is "personal information," the non-disclosure period is set to "indefinite." This indicates that the non-disclosure period is indefinite, and that the non-disclosure period will not become available to all users M as the period elapses. Furthermore, when the tag information is "tasks," the non-disclosure period is set to "none." This indicates that there is no limit on the period during which the text data and audio data of the audio section cannot be viewed. Furthermore, when the tag information is "todo," the non-disclosure period is set to "one year." This indicates that the non-disclosure period of one year will become available to all users M after the period has elapsed.

[0070] Although the present embodiment has been described above, the present invention is not limited to the above-described embodiment. Furthermore, the effects of the present invention are not limited to those described in the above-described embodiment. For example, the system configuration shown in FIG. 1 and the hardware configurations shown in FIGS. 2 and 3 are merely examples for achieving the object of the present invention and are not particularly limited. Furthermore, the functional configuration shown in FIG. 4 is also merely an example and is not particularly limited. It is sufficient for the information processing system 1 of FIG. 1 to be provided with the functionality to execute the above-described processes as a whole, and the functional configuration used to realize these functions is not limited to the example of FIG. 4. For example, in the above-described embodiment, the management server 10 performs the process of decrypting encrypted text data and audio data, but the decryption process may also be performed by the client terminal 30.

[0071] 5 and 6 are merely examples and are not particularly limited. The steps are not necessarily performed in chronological order according to the illustrated order of steps, but may be performed in parallel or individually. The screens shown in FIGS. 10 to 14 are also merely examples and are not particularly limited. Any type of user interface that allows text data and audio data to be selected and viewed in units of audio segments can be displayed on the client terminal 30.

[0072] In the above-described embodiment, three levels of security, "high," "medium," and "low," are used as the security levels assigned to tag information of speech segments. However, this is merely an example, and any level that can indicate the degree of confidentiality can be used. For example, three levels, "A" to "C," or five levels, "1" to "5," may be used. In the above-described embodiment, one viewing authority level is assigned to one user M, but this is not limiting. For example, a security level may be assigned to each department depending on the department to which user M belongs or the role of user M. In this case, the security level for each department may be displayed on the screen on the right side of FIG. 10.

[0073] Furthermore, in the above-described embodiment, encryption and decryption are used to control the output of text data and audio data in the speech section, but this is merely an example. Any manner that allows control over whether the text data and audio data can be viewed can be employed. For example, even without encryption and decryption, whether the text data and audio data can be viewed can be controlled by not displaying the audio data play button in the user interface as described above, making the play button unpressable, or blacking out or masking the relevant parts of the text data.

[0074] In the above embodiment, a security level is assigned to each speech segment, but this is not limiting. For example, a security level may be assigned to each morpheme that constitutes a speech segment.

[0075] Furthermore, the tag information shown in Figure 8 and the like is merely an example. For example, the security level assigned to the tag information "Meeting Information" is set to "Low," but this can also be divided into "Meeting Information (1)" for executive-level meetings, "Meeting Information (2)" for department-level meetings, and "Meeting Information (3)" for team meetings. In this case, the security level can be hierarchically set as "High" for "Meeting Information (1)," "Medium" for "Meeting Information (2)," and "Low" for "Meeting Information (3)." [Explanation of symbols]

[0076] 1...information processing system, 10...management server, 11...control unit, 30...client terminal, 31...control unit, 90...network, 101...voice data acquisition unit, 102...text conversion unit, 103...text data classification unit, 104...voice data classification unit, 105...level assignment unit, 106...encryption unit, 107...access acceptance unit, 108...decryption unit, 109...output control unit

Claims

1. a processor; The processor: Dividing the voice data into text data and the voice data into a plurality of voice segments; assigning a security level to each of the plurality of voice segments according to the contents of the text data and voice data of each of the plurality of voice segments; In accordance with the security level and the level of authority associated with the accessing user, processing is performed to restrict the user's viewing of the text data and listening to the audio data for each of the plurality of audio segments selected from the list, and output of the text data and audio data for the selected audio segment is controlled. Information processing device.

2. the processor assigns the security level to each of the plurality of speech segments according to the type of each of the plurality of morphemes constituting each of the plurality of speech segments, as the content. The information processing device according to claim 1 .

3. the processor controls the output of at least one of each of the plurality of speech segments and each of the plurality of morphemes according to the security level. The information processing device according to claim 2 .

4. The processor controls the output by outputting the portion of the text data and the speech data containing the morpheme in a predetermined output format. The information processing device according to claim 3 .

5. When the type of the morpheme is a proper noun or a numerical value indicating a predetermined content, the processor assigns a higher security level to a speech section including the morpheme than to other speech sections. The information processing device according to claim 2 .

6. the processor assigns the security level, which is preset for each type of content, to each of the plurality of voice segments. The information processing device according to claim 1 .

7. The processor is characterized in that it presets the security level in tag information indicating the type of the content. The information processing device according to claim 6 .

8. the processor divides the text data and the speech data into the plurality of speech segments by at least one of machine learning and natural language processing. The information processing device according to claim 1 .

9. the processor further divides the voice data into the plurality of voice segments by estimating a transition of a subject uttering the voice based on information on changes in voice characteristics obtained from the voice data. The information processing device according to claim 8 .

10. The processor: As a control of the output, a decryption key for decrypting the encrypted audio section to make it viewable is associated with each of the security levels; The decryption key is assigned to the user according to the level of the authority. The information processing device according to claim 1 .

11. The processor further assigns the decryption key to a user according to a predetermined range of the user. The information processing device according to claim 10.

12. The processor further assigns the decryption key to the user in accordance with at least one of an organization to which the user belongs and a role assigned to the user. The information processing device according to claim 10.

13. The processor is further configured to grant the decryption key to the user in accordance with a predetermined period of time. The information processing device according to claim 10.

14. wherein the processor, when the user to whom the decryption key is assigned accesses, performs control to output at least one of text data and voice data of a voice section corresponding to the decryption key. The information processing device according to claim 10.

15. a division means for dividing the voice data into text data and a plurality of voice segments; an assigning means for assigning a security level to each of the plurality of voice segments in accordance with the contents of the text data and the voice data of each of the plurality of voice segments; an output control means for restricting the user's ability to view the text data and listen to the audio data for each of the plurality of selectable audio segments displayed in a list, in accordance with the security level and the level of authority associated with the accessing user, and for controlling the output of the text data and audio data for the selected audio segment; An information processing system comprising:

16. On the computer, A function for converting voice data into text data and dividing the voice data into multiple voice segments; a function of assigning a security level to each of the plurality of voice segments according to the contents of the text data and the voice data of each of the plurality of voice segments; a function of restricting the user's ability to view the text data and listen to the audio data for each of the plurality of selectable audio segments in accordance with the security level and the level of authority associated with the accessing user, and controlling the output of the text data and audio data for the selected audio segment; A program to achieve this.

Citation Information

Patent Citations

  • Voice recording device

    JP2007329794A

  • Selective Security Masking in Recorded Speech Using Speech Recognition Technology

    JP2009501942A

  • Personal information deletion device, method thereof, program thereof, and recording medium

    JP2010271751A

  • Information processing apparatus

    JP2012170024A

  • Information processing device, information processing method and program

    JP2020149628A