Voice actor search system, server, voice actor search program, and voice actor search method
The voice actor search system identifies voice actors in multimedia content by converting audio to text, using dynamic programming and similarity scoring, addressing the inability of existing systems to recognize voice actors and enhancing user engagement.
Patent Information
- Application Number
- JP2021025476
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-02-19
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2041-02-19
AI Technical Summary
Existing systems are unable to identify voice actors in multimedia content based on audio data, despite being able to recognize songs and media streams.
A voice actor search system that includes a user terminal and server, utilizing voice acquisition, conversion to text data, dynamic programming for alignment, and similarity scoring to identify voice actors based on line features and voice characteristics, presenting related content.
Enables accurate identification and presentation of voice actors in multimedia content, enhancing user engagement and expanding the content market by providing information on voice actors.
Smart Images

Figure 0007720152000001 
Figure 0007720152000002 
Figure 0007720152000003
Abstract
Description
[Technical Field]
[0001] The present invention particularly relates to a voice actor search system, a server, a voice actor search program, and a voice actor search method for searching for voice actors based on audio data of multimedia content. [Background technology]
[0002] Conventionally, there have been various systems for identifying multimedia content such as audio or video. Patent Document 1 discloses a conventional system in which a client device continuously performs content identification of a content of a data stream based on a content pattern stored in the client device and queries a server to determine the identity of the content of the data stream. In this case, the client device can capture a media stream (e.g., radio) and request the server to search a database of media records (also known as media tracks) for matches to identify the media stream. The technology in Patent Document 1 identifies songs from audio stream data. Specifically, it searches for the song title itself by searching the "fingerprint," "sub-fingerprint," and "watermark" of the stream. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2014-516189 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the system described in Patent Document 1 was unable to search for who was playing the characters included in the content, that is, who the voice actors were.
[0005] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a voice actor search system that solves the above-mentioned problems. [Means for solving the problem]
[0006] The voice actor search system of the present invention is a voice actor search system comprising a user terminal and a server connected to the terminal, wherein the terminal comprises a voice acquisition unit that acquires voice data containing the voice of a voice actor, and a voice transmission unit that transmits the voice data acquired by the voice acquisition unit to the server, and the server comprises a voice receiving unit that receives the voice data from the terminal, a voice conversion unit that converts the voice data received by the voice receiving unit into text data, a voice actor search unit that searches for lines of the voice actor and identifies the voice actor based on the text data converted by the voice conversion unit, and a content presentation unit that presents content related to the voice actor identified by the voice actor search unit to the terminal. a line acquisition unit that acquires the lines of the voice actors from the network and stores them as a database; The voice actor search unit comprises: The aforementioned Based on the lines of the voice actor stored in the database, the voice actor is identified, and the lines of the voice actor are aligned using DP (Dynamic Programming) taking into account conversion errors and differences in the data in the database, rather than being used as is, and a similarity score is calculated. Analysis The system calculates a similarity score based on the feature model of the linguistic sequence of the dialogue itself, and outputs the lines with the highest calculated similarity scores as search results. The line acquisition unit calculates line features included in the lines of the voice actor from the database, and the line features indicate line features that indicate the characteristics of the voice actor even if the content is different, including conventions about the voice actor's role and requested lines, and the voice actor search unit limits a search space based on the line features for lines included in the text data, or performs a search based on the line features, and identifies the voice actor. It is characterized by: The voice actor search system of the present invention is characterized in that the line acquisition unit calculates the voice actor's line habits, including the phrasing, endings, and tone of voice, as line features using a natural language analysis AI. In the voice actor search system of the present invention, the voice conversion unit acquires voice features from the voice data and includes these voice features in the text data, and the voice features include, in addition to features used in voiceprint recognition, vocalization features, voice quality, intonation of the lines, and pronunciation features due to dialects and accents, and the voice quality includes features such as the gender of the voice actor, estimated age, and whether the voice is energetic or soothing, and the voice actor search unit searches for the lines of the voice actor based on the voice features included in the text data. The server of the present invention is a server for a voice actor search system, and includes a voice receiving unit that receives voice data containing the voice of a voice actor from a user's terminal, a voice conversion unit that converts the voice data received by the voice receiving unit into text data, a voice actor search unit that searches for lines of the voice actor based on the text data converted by the voice conversion unit to identify the voice actor, and a content presentation unit that presents content related to the voice actor identified by the voice actor search unit to the terminal. a line acquisition unit that acquires the lines of the voice actors from the network and stores them as a database; The voice actor search unit comprises: The aforementioned Based on the lines of the voice actor stored in the database, the voice actor is identified, and the lines of the voice actor are aligned using DP (Dynamic Programming) taking into account conversion errors and differences in the data in the database, rather than being used as is, and a similarity score is calculated. Analysis The system calculates a similarity score based on the feature model of the linguistic sequence of the dialogue itself, and outputs the lines with the highest calculated similarity scores as search results. The line acquisition unit calculates line features included in the lines of the voice actor from the database, and the line features indicate line features that indicate the characteristics of the voice actor even if the content is different, including conventions about the voice actor's role and requested lines, and the voice actor search unit limits a search space based on the line features for lines included in the text data, or performs a search based on the line features, and identifies the voice actor. It is characterized by: The voice actor search program of the present invention is a voice actor search program executed by a server of a voice actor search system, which causes the server to receive audio data containing the voice of a voice actor from a user's terminal, convert the received audio data into text data, and search for lines of the voice actor based on the converted text data to identify the voice actor. The lines of the voice actors are acquired from the network and stored in a database,Based on the lines of the voice actor stored in the database, the voice actor is identified, and the lines of the voice actor are aligned using DP (Dynamic Programming) taking into account conversion errors and differences in the data in the database, rather than being used as is, and a similarity score is calculated. Analysis a similarity score based on a feature model of the language sequence in the dialogue itself; and outputting the lines with the highest calculated similarity scores as search results. Calculating line features contained in the lines of the voice actor from the database, the line features indicating line features that indicate the characteristics of the voice actor even if the content is different, including conventions about the voice actor's role and requested lines, and limiting a search space based on the line features for lines contained in the text data or performing a search based on the line features to identify the voice actor; The content relating to the identified voice actor is presented on the terminal. The voice actor search method of the present invention is a voice actor search method executed by a server of a voice actor search system, in which the server receives audio data containing the voice of a voice actor from a user's terminal, converts the received audio data into text data, and searches for lines of the voice actor based on the converted text data to identify the voice actor; The lines of the voice actors are acquired from the network and stored as a database. Based on the lines of the voice actor stored in the database, the voice actor is identified, and the lines of the voice actor are aligned using DP (Dynamic Programming) taking into account conversion errors and differences in the data in the database, rather than being used as is, and a similarity score is calculated. Analysis a similarity score based on a feature model of the linguistic sequence of the dialogue itself; and outputting the lines with the highest calculated similarity scores as search results. Calculating line features contained in the lines of the voice actor from the database, the line features indicating line features that indicate the characteristics of the voice actor even if the content is different, including conventions and requested lines for the voice actor's role, and limiting a search space based on the line features for the lines included in the text data or performing a search based on the line features to identify the voice actor; The method is characterized in that content related to the identified voice actor is presented to the terminal. [Effects of the Invention]
[0007] According to the present invention, a voice actor search system can be provided that can search for voice actors by converting audio data received from a terminal into text data, searching for the voice actor's lines based on this text data to identify the voice actor, and presenting content related to the identified voice actor to the terminal. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a system configuration diagram of a voice actor search system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing a control configuration of the server shown in FIG. [Figure 3] 1 is a block diagram showing the functional configuration of a voice actor search system according to an embodiment of the present invention; [Figure 4] FIG. 2 is a block diagram showing a detailed configuration of a voice actor-related DB shown in FIG. [Figure 5] 10 is a flowchart of a voice actor search process according to an embodiment of the present invention. [Figure 6] 6 is a screen example of the voice actor search process shown in FIG. 5. [Figure 7] 10 is a flowchart of a dialogue acquisition process according to an embodiment of the present invention. [Figure 8] 8 is a screen example of the dialogue acquisition process shown in FIG. 7. DETAILED DESCRIPTION OF THE INVENTION
[0009] <Embodiment> [System configuration of voice actor search system X] First, with reference to FIG. 1, the system configuration of a voice actor search system X according to an embodiment of the present invention will be described. The voice actor search system X is configured by connecting a server 1 and a user terminal 2 to a network 5 such as the Internet or an intranet.
[0010] The server 1 is a PC (Personal Computer) server, a general-purpose machine, or the like. The server 1 receives audio data 300 (FIG. 3) of multimedia content (hereinafter simply referred to as "content") such as animation or movies as a query from each terminal 2, searches for voice actors from this audio data 300, and presents related content. The server 1 may be a server on a so-called "cloud." The server 1 may also be a server within an intranet.
[0011] Terminal 2 is a computer such as a mobile phone, smartphone, tablet, portable PC, desktop PC, game device, or home appliance used by a user who wants to search for a voice actor.
[0012] Next, the control configuration of the server 1 will be described with reference to FIG. The server 1 includes a control unit 10, a storage unit 11, a network transmission / reception unit 15, etc. Each unit is connected to the control unit 10, and the operation thereof is controlled by the control unit 10.
[0013] The control unit 10 is an information processing unit such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), DSP (Digital Signal Processor), GPU (Graphics Processing Unit), or ASIC (Application Specific Integrated Circuit). The control unit 10 reads out a control program stored in the ROM or HDD of the storage unit 11, expands the control program into RAM, and executes it to operate as each part of the functional blocks described below. The control unit 10 can also set and control the entire device in accordance with predetermined instruction information input from an external management terminal or console (not shown).
[0014] The network transceiver unit 15 is a network connection unit including a LAN board, a wireless transceiver, etc. for connecting to an external network 5. The network 5 in this embodiment is connected via, for example, a LAN, a wireless LAN, a WAN, a mobile phone network, a voice telephone network, etc. The network transmission / reception unit 15 may be capable of transmitting and receiving data over a data communication line and transmitting and receiving voice signals over a voice telephone line.
[0015] The storage unit 11 is a non-transitory recording medium such as a semiconductor memory such as a ROM (Read Only Memory) or a RAM (Random Access Memory) or a HDD (Hard Disk Drive). The ROM and HDD of the storage unit 11 store control programs for controlling the operation of the voice actor search system X, various databases, content data, etc. These control programs include firmware, an OS (Operating System), services and daemons that run on the OS, device drivers and middleware, server software, and application software (hereinafter simply referred to as "apps"). These services (daemons) may include a WWW (World Wide Web) server, CGI (Common Gateway Interface), ASP, server-side programs, etc. In addition, the storage unit 11 may also store user account settings.
[0016] In the voice actor search system X, the control unit 10 may be integrally formed as a CPU with a built-in GPU, a chip-on-module package, a SOC (System On a Chip), etc. The control unit 10 may also include built-in RAM, ROM, flash memory, etc.
[0017] The terminal 2 also has an audio input unit including a microphone and an A / D (Analog to Digital) converter, a network connection unit for transmitting the audio input through the audio input unit to the network 5, a control unit such as a CPU for overall control, and a memory unit for storing a control program.
[0018] [Functional configuration of voice actor search system X] Here, the functional configuration of the voice actor search system X will be described with reference to FIGS. The control unit 10 of the server 1 includes a voice receiving unit 100, a voice converting unit 110, a voice actor searching unit 120, a content presenting unit 130, and a dialogue acquiring unit 140. The storage unit 11 stores voice data 300, text data 310, and a voice actor related DB 320. The terminal 2 includes a voice acquisition unit 200 and a voice transmission unit 210.
[0019] The audio receiving unit 100 receives audio data 300 from the terminal 2. The audio receiving unit 100 acquires the audio data 300 transmitted by the terminal 2 via the network transmitting / receiving unit 15 and stores it in a storage unit. At this time, the audio receiving unit 100 may decode the audio data 300 if it has been encoded.
[0020] The voice conversion unit 110 converts the voice data 300 received by the voice receiving unit 100 into text data 310. The voice conversion unit 110 extracts human voice data by performing frequency analysis using FFT (Fast Fourier Transform) or the like and various filtering on the acquired voice data 300. Then, the voice conversion unit 110 performs voice recognition using various AI (Artificial Intelligence) such as a convolutional NN (Neural Network), a hidden Markov model, or various other voice recognition methods.
[0021] During this speech recognition, the speech conversion unit 110 may acquire speech features of the speech data 300 and include the speech features in the text data 310. The speech features may include, for example, various features used in voiceprint recognition, vocalization features, voice quality, intonation of lines, pronunciation features due to dialects or accents, etc. Among these, the voice quality may be capable of detecting features such as the gender and estimated age of the voice actor, whether the voice is a clear, energetic voice with strong intonation and high tension, or a calm, soothing voice. In this case, the speech conversion unit 110 may also be capable of including the speech features of the speech data 300 in the text data 310.
[0022] The voice actor search unit 120 searches for the lines of a voice actor based on the text data 310 converted by the voice conversion unit 110 to identify the voice actor. Specifically, the voice actor search unit 120 can identify the voice actor based on the lines of the voice actor stored in the voice actor-related DB 320. In this case, the voice actor search unit 120 may also be able to identify the voice actor using voice features. Furthermore, the voice actor search unit 120 may also be able to identify the voice actor based on line features calculated by the line acquisition unit 140, which will be described below.
[0023] The content presentation unit 130 presents to the terminal 2 content related to the voice actor identified by the voice actor search unit 120. The content presentation unit 130 can, for example, refer to the voice actor related DB 320 and present to the terminal 2 information about the voice actor identified as a search result for the voice actor, and information related to this voice actor that can be effectively utilized by the user.
[0024] The line acquisition unit 140 acquires lines of voice actors from the network 5 and stores them in the voice actor-related DB 320. For example, the line acquisition unit 140 stores memorable lines of voice actors entered by users on a dedicated website in line data 400 ( FIG. 4 ) of the voice actor-related DB 320. Alternatively, the line acquisition unit 140 may automatically and periodically acquire lines posted on the voice actor's website via the network 5 and store them in the voice actor-related DB 320. The line acquisition unit 140 may perform this automatic and periodic acquisition using a program such as a crawler that acquires data including HTML (Hyper Text Markup Language) and content from a WWW site using the HTTP protocol, such as wget. Furthermore, the line acquisition unit 140 may perform voice recognition on the audio data of the acquired content and store the recognition results in the line data 400. Then, the line acquisition unit 140 may calculate line feature amounts included in the lines of the voice actors stored in the voice actor-related DB 320. In this case, the line acquisition unit 140 may calculate line features using, for example, a natural language processing AI, and may be able to calculate features for each voice actor.
[0025] The audio acquisition unit 200 acquires audio data 300 containing the voice of a voice actor. In this embodiment, while the user is viewing content, the user issues an instruction using a web browser or a dedicated app on the terminal 2. In response, an audio signal is input as audio data 300 from the audio input unit, and the audio acquisition unit 200 acquires this.
[0026] The audio transmitting unit 210 transmits the audio data 300 acquired by the audio acquiring unit 200 to the server 1. The audio transmitting unit 210 transmits the audio data 300 to the server 1 via the network connection unit and over the network 5. At this time, the audio transmitting unit 210 can encode the audio data 300 using various codecs and transmit the encoded audio data 300. The audio transmitting unit 210 may encrypt the audio data 300 when encoding it.
[0027] The voice data 300 is data of voice acquired from the terminal 2. This voice data 300 may be coded using various codecs or encrypted when acquired. Furthermore, the voice data 300 may include only signals in the frequency band of human voice required for voice recognition.
[0028] The text data 310 is text (character) data obtained by speech recognition from the speech data 300. In this embodiment, the text data 310 mainly includes text of lines spoken by a voice actor. Furthermore, the text data 310 may also include, as text data, speech features of the speech data 300 when speech conversion is performed by the speech conversion unit 110.
[0029] The voice actor related DB 320 is a database of lines of voice actors and related content that is referenced when identifying voice actors from the text data 310. The detailed configuration of the voice actor related DB 320 will be described later.
[0030] Next, the voice actor-related DB 320 will be described in detail with reference to FIG. The voice actor related DB 320 includes line data 400 , voice feature data 410 , line feature data 420 , and related content data 430 .
[0031] The dialogue data 400 is a database containing text data of dialogues by voice actors. The dialogue data 400 may be configured as a table containing various data associated with the dialogues, such as the type of content, the name, and the number of episodes of the program.
[0032] The audio feature data 410 is data indicating audio features for each voice actor. This audio feature data 410 may be an accumulation of audio features obtained by analyzing the audio data 300 transmitted from the terminal 2. Alternatively, the audio feature data 410 may be data of audio features based on the audio data 300 acquired by a crawler. This audio feature data 410 may also be stored in text format.
[0033] The line feature data 420 is a database that stores line feature amounts calculated by natural language processing. This line feature data 420 may be, for example, a script written to indicate the characteristics of a voice actor across different works, or may be characteristics based on the ad-lib habits of each voice actor.
[0034] The related content data 430 includes, for example, data on the voice actor's website, home page, SNS (Social Network Service) account, etc., and related content introduction and sales sites, etc. In addition, the related content data 430 may also include files of the content itself, files of the voice actor's voice data, etc.
[0035] In addition, the voice actor related DB 320 may be able to store data on the content of each voice actor.
[0036] Here, the control unit 10 of the server 1 of the voice actor search system X executes a control program stored in the storage unit 11, thereby functioning as a voice receiving unit 100, a voice conversion unit 110, a voice actor search unit 120, a content presenting unit 130, and a dialogue acquiring unit 140. These may be provided as services (daemons). Furthermore, the control unit of the terminal 2 is made to function as a voice acquisition unit 200 and a voice transmission unit 210 by executing a control program stored in the storage unit. Furthermore, each part of the server 1 and terminal 2 of the voice actor search system X described above is a hardware resource for executing the voice actor search method of the present invention. Note that a part or any combination of the above-described functional configurations may be configured in terms of hardware or circuits using ICs, programmable logic, FPGAs (Field-Programmable Gate Arrays), or the like.
[0037] [Voice actor search processing using voice actor search system X] Next, the voice actor search process performed by the voice actor search system X according to the embodiment of the present invention will be described with reference to FIGS. In the voice actor search process according to this embodiment, the terminal 2 acquires audio data 300 containing the voice of the voice actor. The audio data 300 acquired by the audio acquisition unit 200 is then transmitted to the server 1. The server 1 receives the audio data 300 containing the voice of the voice actor from the user's terminal. The received audio data 300 is then converted into text data 310. The voice actor's lines are then searched for and the voice actor is identified based on the converted text data 310. Content related to the identified voice actor is then presented to the terminal 2.
[0038] In the voice actor search process of this embodiment, the control unit 10 of the server 1 executes the control program stored in the memory unit 11, and the control unit of the terminal 2 executes the control program stored in the memory unit, each in cooperation with the respective units, using hardware resources. The voice actor search process will be described in detail below for each step with reference to the flowchart in FIG.
[0039] (Step S201) First, the voice acquisition unit 200 of the terminal 2 performs voice data acquisition processing. For example, suppose a user is watching content via television broadcasting or video-on-demand, displaying it on a television screen or the display of terminal 2 (hereinafter referred to as "television, etc.") and playing audio from speakers, etc. At this time, the user may want to know who the voice actors are who appear in the content, and what other works they have appeared in, and may find out about a voice actor search service provided by server 1 on a search site (not shown). The user then accesses the website of server 1's service using a web browser or a dedicated app, etc. (hereinafter simply referred to as "app, etc."). Server 1 then instructs it to transmit audio from the content.
[0040] Here, the user may instruct "record" using an app or the like when a voice actor's lines for the content are being played on a television or the like. Alternatively, the user may instruct "record" using an app or the like for content being played on terminal 2. Upon receiving these instructions, the audio capture unit 200 records the audio played on a television or the like from the audio input unit as audio data 300, or captures a portion of the content data as audio data 300. This allows the audio capture unit 200 to acquire audio data 300 that includes the voice of the voice actor.
[0041] (Step S202) Next, the voice transmission unit 210 performs voice data transmission processing. The voice transmitting unit 210 transmits the voice data 300 acquired by the voice acquiring unit 200 to the server 1. At this time, the voice transmitting unit 210 can filter, compress, encode, encrypt, etc. the voice data 300 to reduce it to only signals in the frequency band of human voice necessary for voice recognition, and then transmit the data.
[0042] (Step S101) Here, the voice receiving unit 100 of the server 1 performs voice data receiving processing. The voice receiving unit 100 receives the voice data 300 from the terminal 2 and temporarily stores it in the storage unit 11. At this time, the voice receiving unit 100 may also perform decoding and the like.
[0043] (Step S102) Next, the speech conversion unit 110 performs speech-to-text conversion processing. The speech conversion unit 110 converts the received speech data 300 into text data 310. The speech conversion unit 110 is capable of performing speech recognition using various speech recognition methods. Here, if the speech conversion unit 110 is unable to recognize meaningful words as dialogue or if the likelihood or accuracy of recognition is low, it may convert the dialogue to an appropriate sentence by adding words that are as close as possible or by supplementing the context.
[0044] (Step S103) Next, the speech conversion unit 110 performs a feature addition process. The speech conversion unit 110 can also calculate speech features of the speech data 300. Then, the speech conversion unit 110 includes the calculated speech features in the text data 310. For example, the speech conversion unit 110 can add, in addition to voiceprint recognition data, feature information such as "intonation: Kansai, Hokuriku" and "feature: female, 20s to 30s, soothing voice" as speech features.
[0045] (Step S104) Next, the voice actor search unit 120 performs a dialogue text voice actor identification process. The voice actor search unit 120 searches the voice actor related DB 320 for lines based on the text data 310 converted by the voice conversion unit 110, and identifies the voice actor based on the lines.
[0046] Specifically, the voice actor search unit 120 first searches for lines included in the text data 310 based on the line data 400 stored in the voice actor-related DB 320. In this case, the voice actor search unit 120 may search for similar lines by taking into account conversion errors and differences in data within the database, rather than simply searching for the lines themselves. More specifically, the voice actor search unit 120 may align lines using, for example, dynamic programming (DP) or the like, and calculate a similarity score. Alternatively, the voice actor search unit 120 may calculate a similarity score using, for example, morphological analysis of natural language analysis, and search for similar lines. Furthermore, the voice actor search unit 120 may also calculate a similarity score using, for example, a feature model of the arrangement of words in the lines themselves. The voice actor search unit 120 may then output, as search results, several lines with the highest calculated similarity scores. More specifically, for example, the voice actor search unit 120 may be able to search for a similar line, "I thought something like this might happen," stored in the line data 400, from the line, "I thought something like this might happen," in the text data 310.
[0047] Furthermore, the voice actor search unit 120 may be able to search the dialogue data 400 based on the feature information of the voice features included in the text data 310, taking into consideration the voice feature data 410. More specifically, in the above example, the voice search unit may narrow the search space and search for voice actors who fit the criteria of "intonation: Kansai, Hokuriku" and "feature: female, 20s to 30s, soothing voice." In this case, if no lines with a similarity score higher than a predetermined threshold are detected from the dialogue data 400, the voice actor's representative lines that match the feature information may be searched for from the dialogue data 400.
[0048] Additionally, the voice actor search unit 120 may limit the search space for lines included in the text data 310 based on the line feature data 420. Alternatively, the voice actor search unit 120 may also be able to perform a search based on this line feature data 420. In this case, if no matching lines are detected in the line data 400, the voice actor search unit 120 may search for representative lines of the voice actor that match the line feature data 420 from the line data 400.
[0049] The voice actor search unit 120 identifies the voice actor who spoke the lines from these retrieved lines. At this time, the voice actor search unit 120 does not need to include voice actors who appear anonymously (under a different name) in the identification results. The voice actor search unit 120 can also store the search results in the voice actor related DB 320.
[0050] (Step S105) Next, the content presenting unit 130 performs a related content presenting process. The content presenting unit 130 acquires data to be presented to the terminal 2 (hereinafter referred to as "search result data") from the related content data 430 for the voice actor identified by the voice actor searching unit 120, and transmits it to the terminal 2. Specifically, for example, the content presenting unit 130 can transmit search result data including HTML or the like based on the data of the voice actor in the related content data 430 to the terminal 2. At this time, the content presenting unit 130 can also add related hyperlinks (hereinafter referred to as "links"), content videos, images, photographs, audio data, and links thereto.
[0051] (Step S203) Next, the control unit and display unit of the terminal 2 perform a result display process. The terminal 2 obtains and presents the search result data from the server 1. Specifically, for example, data of related content related to the identified voice actor is displayed on the display unit or the like as a screen of an application or the like.
[0052] The example screen 500 in FIG. 6 shows an example of how this search result data is displayed. This example shows two candidate voice actors searched for based on the lines and voice feature data 410 included in the text data 310. The example screen 500 displays a photo of the voice actor, the voice actor's name, their agency (agency) and its website, a list of works they have appeared in and links to them, and links to related products. Clicking on the list of works and their links may allow playback of the content of the work via video on demand (VOD). Alternatively, the link may play a sample audio of the voice actor or special lines addressed to the user. Furthermore, this playback may be a link to a video playback site or a link to streaming content stored on the server 1. Furthermore, playback may also be possible for playback of e-books, audio data 300, etc.
[0053] Furthermore, related products may not be the content itself in which the voice actor plays a role, but may instead present links to works in which the voice actor is active under the name of the voice actor (or a unit). Alternatively, it is also possible to present links to sales sites, publishers, production committees, etc. for works in which the voice actor appears. This completes the voice actor search process according to the embodiment of the present invention.
[0054] [Line acquisition process using voice actor search system X] Next, with reference to FIGS. 7 and 8, a description will be given of the line acquisition process performed by the voice actor search system X according to the embodiment of the present invention. In the voice actor search process according to this embodiment, lines are acquired over a network and stored as line data 400 in the voice actor related DB 320. Then, line feature amounts included in the lines of the voice actor are calculated from the voice actor related DB 320.
[0055] In the dialogue acquisition process of this embodiment, the control unit 10 of the server 1 executes a control program stored in the storage unit 11 in cooperation with each unit, using hardware resources. The dialogue acquisition process will be described in detail below for each step with reference to the flowchart in FIG.
[0056] (Step S111) The line acquisition unit 140 of the server 1 performs a line acquisition process. In this process, the line acquisition unit 140 acquires the lines of the voice actor and stores them as line data 400 in the voice actor related DB 320.
[0057] For example, the line acquisition unit 140 can have a user input lines via the terminal 2 connected to the network 5 and acquire the lines. Specifically, a user accesses a website such as "Minna no Taikou" (Everyone's Lines) via an app or the like, which allows the user to input lines for Server 1. Then, the line acquisition unit 140 of the server 1 instructs the user to transmit the lines.
[0058] An example of a form for allowing a user to input lines will be described with reference to screen example 510 in Fig. 8. The line acquisition unit 140 allows the user to input, for example, the name of a voice actor, the title of a work, the name of a character (appearance), lines, characteristics of the lines (explanation, comments), etc. in the form. Furthermore, it may be possible to allow the user to input additional information such as the episode in which the character appears in the work. The information entered using this form or the like is acquired by the line acquisition unit 140 via the network transmission / reception unit 15 and stored in the line data 400 of the voice actor related DB 320.
[0059] Alternatively, the line acquisition unit 140 may periodically acquire lines posted on the voice actor's website via the network 5 using a crawler or the like, and store the lines in the line data 400. This is because the voice actor's website may present lines from content in which the voice actor has appeared.
[0060] Furthermore, the line acquisition unit 140 may identify characters in the content and the voice actors who play them from websites such as online encyclopedias and content descriptions. In this case, the line acquisition unit 140 may acquire memorable lines of these characters from websites such as reviews of the content and social networking sites, and store them in the line data 400. For example, as a specific example, a crawler may be used to identify the voice actor of the villain "Roy" in the work "Herotale of the Galaxy," acquire the lines from a "site that compiles famous quotes and lines" for the work, and store the lines together with the source in the line data 400.
[0061] In these cases, it may be possible to acquire keywords about the voice characteristics of each voice actor from the voice actor's website, SNS, review website, etc., and store them in the voice feature data 410. Specifically, in the above example, the voice actor's gender, age, and characteristics such as "soothing" can be stored in the voice feature data 410.
[0062] Alternatively, the line acquisition unit 140 can read the script data of the content, acquire the lines of each voice actor, and store them in the line data 400. Furthermore, the dialogue acquisition unit 140 can acquire voice data from the content itself, perform voice recognition on this voice data, recognize the lines, and store them in the dialogue data 400. This content loading may also be performed semi-automatically by searching using a crawler. Furthermore, the dialogue acquisition unit 140 may calculate voice feature data from the loaded voice data and store the data in the voice feature data 410.
[0063] (Step S112) The dialogue acquisition unit 140 performs a dialogue feature calculation process. The line acquisition unit 140 calculates line features based on the stored line data 400. This is because, as a voice actor becomes famous, there may be "conventions" and "requested lines" for the character. In other words, the line acquisition unit 140 can extract line features, which are line features that indicate the characteristics of the voice actor, even across different content. Specifically, the line acquisition unit 140 calculates, as line features, the "habits" of the voice actor's lines, such as the voice actor's phrasing, endings, and tone of voice, using natural language analysis AI or the like. This may be lines from the script or lines ad-libbed by the voice actor. Specifically, for example, it may be possible to calculate line features such as the voice actor often playing "pretentious villains" and often saying things like "That's..." In this way, it is possible to use the dialogue feature data 420 to limit the search space as described above, or to perform a search using the dialogue feature. This completes the dialogue acquisition process according to the embodiment of the present invention.
[0064] The above configuration can provide the following effects. In the past, when viewing anime, foreign films, and other content, people have often become interested in the actors in the film. However, when a voice actor, rather than a live-action actor, plays or dubs a character, even if you are interested in the voice actor, you have to check their name in the end credits and then search for related content yourself, which can be a hassle. However, conventional music search systems such as that described in Patent Document 1 cannot search for voice actors.
[0065] In contrast, the voice actor search system X according to an embodiment of the present invention is a voice actor search system comprising a user's terminal 2 and a server 1 connected to the terminal 2, wherein the terminal 2 comprises an audio acquisition unit 200 that acquires audio data 300 containing the voice of the voice actor, and an audio transmission unit 210 that transmits the audio data 300 acquired by the audio acquisition unit 200 to the server 1, and the server 1 comprises an audio receiving unit 100 that receives the audio data 300 from the terminal 2, an audio conversion unit 110 that converts the audio data 300 received by the audio receiving unit 100 into text data 310, a voice actor search unit 120 that searches for the voice actor's lines and identifies the voice actor based on the text data 310 converted by the audio conversion unit 110, and a content presentation unit 130 that presents content relating to the voice actor identified by the voice actor search unit 120 to the terminal 2. By configuring it in this way, it is possible to provide a voice actor search system that allows users to search for voice actors who appear in content. It also makes it possible to further utilize content from the perspective of voice actors. It is also expected to have the effect of increasing the number of voice actor fans. In addition, users will be able to view content related to voice actors that interest them, making it possible to make better use of the content itself. As a result, it is possible to liven up the content market itself.
[0066] In the voice actor search system X according to an embodiment of the present invention, the server 1 further includes a line acquisition unit 140 that acquires lines of a voice actor from the network and stores them as line data 400 in the voice actor related DB 320, and the voice actor search unit 120 identifies the voice actor based on the line data 400 stored in the voice actor related DB 320. By configuring it in this way, lines can be acquired by users, crawlers, etc. and stored in a database. This allows voice actors to be searched for without the administrator of the voice actor search system X having to prepare the entire database. This reduces costs. Also, by users entering lines, the voice actor search site can be made more popular, and advertising revenue can be expected. Furthermore, users can collect memorable lines, which can expand the database. In addition, it is possible to understand the popularity of voice actors and the content itself.
[0067] In the voice actor search system X according to the embodiment of the present invention, the voice conversion unit 110 acquires voice features of the voice data 300 and includes the voice features in the text data 310. With this configuration, voice actor searches can be performed using voice features in addition to information obtained through voice recognition. This increases the accuracy of identifying voice actors. It also makes it possible to identify voice actors who have not yet been registered in the voice actor-related DB 320. Furthermore, by providing a voice actor that matches the features as a search result, even if the identified voice actor is different from the actual voice actor, it is possible to present a voice actor that matches the user's intention. This increases the user's interest in voice actors.
[0068] In the voice actor search system X according to an embodiment of the present invention, the line acquisition unit 140 also calculates the line features contained in the lines of the voice actor from the voice actor related DB 320, and the voice actor search unit 120 identifies the voice actor based on the line features calculated by the line acquisition unit 140. This configuration makes it possible to search and identify voice actors based on the characteristics of the lines themselves. This narrows the search space, making the search more efficient and allowing for a more appropriate voice actor to be found. It is also possible to search for voice actors based on their line habits, the "conventions" and "requested lines" of the role, etc.
[0069] Other Embodiments In the above embodiment, an example has been described in which the server 1 recognizes voice and converts it into text to perform a search. However, the lines may be converted into text data 310 in an app or the like on terminal 2, or may be converted into text data 310 by another so-called cloud-based service. Furthermore, the voice actor-related DB 320 may be provided on server 1, and searches may be performed by an app or the like on terminal 2. Alternatively, the voice actor-related DB 320 may be provided as data in an app or the like on terminal 2, so that voice actors can be searched for solely on terminal 2. Furthermore, the app or the like may be configured as a normal installable app, a web app, or a temporary app that accesses the website of server 1 from a web browser. This configuration allows for flexible adaptation to various configurations.
[0070] In the above embodiment, the configuration in which the voice data 300 is transmitted to the server 1 has been described. However, it may also be possible for the user to directly input the text data 310 of the lines and send it to the server 1 to search for the voice actor. By configuring it in this way, it becomes possible to directly search for voice actors based on popular lines and slang on network 5.
[0071] In the above embodiment, an example has been described in which the speech feature quantity of the speech data 300 converted by the speech conversion unit 110 is also added to the text data 310 as text. However, the speech feature may be used for searching as numerical data as it is. In this case, the speech feature data 410 may also contain numerical data. This configuration makes it possible to efficiently search using voice features. Furthermore, features that cannot be converted into text can also be used to identify voice actors.
[0072] Furthermore, it goes without saying that the configurations and operations of the above-described embodiments are merely examples, and can be modified as appropriate within the scope of the present invention. [Explanation of symbols]
[0073] 1 Server 2. Devices 5. Network 10 Control Unit 11 Storage section 15 Network Transmitter / Receiver 100 Audio receiving unit 110 Voice conversion unit 120 Voice Actor Search Department 130 Content presentation section 140 Line acquisition part 200 Audio acquisition unit 210 Audio transmission unit 300 audio data 310 Text Data 320 Voice Actor Related DB 400 dialogue data 410 Audio feature data 420 Dialogue Feature Data 430 Related Content Data 500, 510 screen example X Voice Actor Search System
Claims
1. A voice actor search system comprising a user terminal and a server connected to the terminal, The terminal a voice acquisition unit that acquires voice data containing the voice of a voice actor; a voice transmission unit that transmits the voice data acquired by the voice acquisition unit to the server, The server a voice receiving unit that receives the voice data from the terminal; a voice conversion unit that converts the voice data received by the voice receiving unit into text data; a voice actor search unit that searches for lines of the voice actor based on the text data converted by the voice conversion unit and identifies the voice actor; a content presentation unit that presents content related to the voice actor identified by the voice actor search unit to the terminal; a line acquisition unit that acquires the lines of the voice actors from a network and stores them as a database; the voice actor search unit identifies the voice actor based on the lines of the voice actor stored in the database; Instead of using the lines of the voice actor as is, the lines of the voice actor are aligned using DP (Dynamic Programming) to calculate a similarity score, taking into account conversion errors and differences in the data in the database, a similarity score is calculated using morphological analysis of natural language analysis, a similarity score is calculated using a feature model of the arrangement of words in the lines themselves, and several lines with the highest calculated similarity scores are output as search results, the line acquisition unit calculates line features included in the lines of the voice actor from the database; The line feature amount indicates line features that indicate the characteristics of the voice actor even if the content is different, including conventions about the role of the voice actor and requested lines, The voice actor search unit limits a search space based on the line feature amount for lines included in the text data, or performs a search based on the line feature amount, and identifies the voice actor. A voice actor search system characterized by:
2. The line acquisition unit Using AI for natural language analysis, the voice actor's line habits, including the phrasing, endings, and tone of voice, are calculated as line features.
2. The voice actor search system according to claim 1.
3. the speech conversion unit acquires speech features of the speech data and includes the speech features in the text data; The voice features include features used for voiceprint recognition, as well as vocalization features, voice quality, intonation of lines, and pronunciation features due to dialects and accents. The voice quality includes features such as the gender of the voice actor, estimated age, and whether the voice is energetic or soothing. The voice actor search unit searches for lines of the voice actor based on the voice feature amount included in the text data.
3. The voice actor search system according to claim 1 or 2.
4. A server for a voice actor search system, an audio receiving unit that receives audio data containing the voice of a voice actor from a user's terminal; a voice conversion unit that converts the voice data received by the voice receiving unit into text data; a voice actor search unit that searches for lines of the voice actor based on the text data converted by the voice conversion unit and identifies the voice actor; a content presentation unit that presents content related to the voice actor identified by the voice actor search unit to the terminal; a line acquisition unit that acquires the lines of the voice actors from a network and stores them as a database; the voice actor search unit identifies the voice actor based on the lines of the voice actor stored in the database; Instead of using the lines of the voice actor as is, the lines of the voice actor are aligned using DP (Dynamic Programming) to calculate a similarity score, taking into account conversion errors and differences in the data in the database, a similarity score is calculated using morphological analysis of natural language analysis, a similarity score is calculated using a feature model of the arrangement of words in the lines themselves, and several lines with the highest calculated similarity scores are output as search results, the line acquisition unit calculates line features included in the lines of the voice actor from the database; The line feature amount indicates line features that indicate the characteristics of the voice actor even if the content is different, including conventions about the role of the voice actor and requested lines, The voice actor search unit limits a search space based on the line feature amount for lines included in the text data, or performs a search based on the line feature amount, and identifies the voice actor. A server characterized by:
5. A voice actor search program executed by a server of a voice actor search system, The server, The user's device receives audio data containing the voice of the voice actor, converting the received voice data into text data; searching for lines of the voice actor based on the converted text data to identify the voice actor; The lines of the voice actor are acquired from the network and stored in a database. Identifying the voice actor based on the lines of the voice actor stored in the database; Instead of using the lines of the voice actor as they are, the lines of the voice actor are aligned using DP (Dynamic Programming) to calculate a similarity score, taking into consideration conversion errors and differences in the data in the database, a similarity score is calculated using morphological analysis of natural language analysis, a similarity score is calculated using a feature model of the arrangement of words in the lines themselves, and several lines with the highest calculated similarity scores are output as search results, Calculating line features included in the lines of the voice actor from the database; The line feature amount indicates line features that indicate the characteristics of the voice actor even if the content is different, including conventions about the role of the voice actor and requested lines, For lines included in the text data, a search space is limited based on the line feature amount, or a search is performed based on the line feature amount, and the voice actor is identified; and presenting content relating to the identified voice actor on the terminal. A voice actor search program characterized by:
6. A voice actor search method executed by a server of a voice actor search system, The server Receives audio data containing the voice of the voice actor from the user's device, converting the received voice data into text data; searching for lines of the voice actor based on the converted text data to identify the voice actor; The lines of the voice actors are acquired from the network and stored as a database. Identifying the voice actor based on the lines of the voice actor stored in the database; Instead of using the lines of the voice actor as is, the lines of the voice actor are aligned using DP (Dynamic Programming) to calculate a similarity score, taking into account conversion errors and differences in the data in the database, a similarity score is calculated using morphological analysis of natural language analysis, a similarity score is calculated using a feature model of the arrangement of words in the lines themselves, and several lines with the highest calculated similarity scores are output as search results, Calculating line features included in the lines of the voice actor from the database; The line feature amount indicates line features that indicate the characteristics of the voice actor even if the content is different, including conventions about the role of the voice actor and requested lines, For lines included in the text data, limiting a search space based on the line feature amount or performing a search based on the line feature amount to identify the voice actor; Presenting content related to the identified voice actor to the terminal. A voice actor search method characterized by:
Citation Information
Patent Citations
Merchandise information providing system
JP2002073665A
Virtual shopping mall system, associating device, and program
JP2011141848A
Detail information management system
JP2012205186A
Display device, interactive server and response information provision method
JP2014003610A
A method and system for performing a comparison of received data and providing subsequent services based on that comparison.
JP2014516189A