Information processing device, information processing method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TOKYO BROADCASTING SYSTEM TELEVISION INC
- Filing Date
- 2025-05-30
- Publication Date
- 2026-08-03
AI Technical Summary
【0022】 本発明によれば、媒体に含まれる文字情報の不備を、媒体の種別に応じて適切に判定することが可能となる。
Smart Images

Figure 0007899402000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Conventionally, in the field of broadcasting, program production and broadcasting using media such as manuscripts and flip teleprompts have been carried out. It is extremely important to check whether there are any deficiencies such as misspellings in the character information contained in such media. In this regard, systems for detecting typographical errors and inconsistencies (contradictions and notation variations) in manuscripts have been utilized. For example, Patent Document 1 describes a system for extracting typographical error information from a document based on pre-stored typographical error patterns.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the field of broadcasting, the genres and contents handled by programs are diverse, and accordingly, the criteria for determining whether specific character information contained in the media is defective may also vary. Therefore, the burden of determining deficiencies in the character information of the media has been large.
[0005] Therefore, an object of the present invention is to provide an information processing apparatus, an information processing method, and a program capable of appropriately determining deficiencies in character information contained in a media according to the type of the media.
Means for Solving the Problems
[0006] An information processing device according to one aspect of the present invention comprises: a media data acquisition unit that acquires media data; a first type determination unit that determines a first type, which is the type of media data, based on the media data; an extraction unit that extracts at least one candidate for deficiencies in character information contained in the media data; a determination unit that determines deficiencies in character information based on the first type and at least one candidate; and an output unit that outputs the determination result from the determination unit.
[0007] According to this embodiment, a determination of character information defects is made based on a first type, which is the type of media data, and at least one candidate for defects in character information contained in the media data, and the determination result is output. Therefore, it becomes possible to appropriately determine character information defects contained in the media according to the type of media.
[0008] In the embodiment described above, the extraction unit may include at least one of the following: a first extraction unit that extracts a first candidate from at least one candidate by referring to a predetermined typographical error database, and a second extraction unit that extracts a second candidate from at least one candidate by inputting a predetermined prompt (an instruction sentence to be input to the language model) to a predetermined language model.
[0009] According to this embodiment, the first candidate extracted by referring to a predetermined typographical error database, the second candidate extracted by inputting a predetermined prompt into a predetermined language model, and at least one candidate for character information defects contained in the media data can be included, making it possible to extract candidates for character information defects more comprehensively and with greater accuracy.
[0010] In the embodiment described above, the extraction unit may further include a typo database determination unit that determines a predetermined typo database from a plurality of typo databases based on a first type.
[0011] According to this embodiment, a typographical error database suitable for the first type can be used as the predetermined typographical error database referenced for extracting the first candidate, thereby improving the accuracy of extracting candidates for deficiencies in character information.
[0012] In the embodiments described above, the extraction unit may further include a prompt determination unit that determines a predetermined prompt from a plurality of prompts that can be input to a predetermined language model based on a first type.
[0013] According to this embodiment, a prompt suitable for the first type can be determined from among multiple prompts as a predetermined prompt to be input to a predetermined language model for extracting a second candidate, thereby improving the accuracy of extracting candidates with incomplete character information.
[0014] In the embodiment described above, the second extraction unit may extract a second candidate associated with a second type, which is the type of deficiency in the character information.
[0015] According to this embodiment, a second candidate corresponding to a second type, which is the type of character information, can be extracted as a second candidate for deficiencies in character information, thereby improving the accuracy of extracting candidates for deficiencies in character information.
[0016] In the embodiment described above, the determination unit may use different criteria for determining whether or not the character information is incomplete, based on the combination of the first type and the second type.
[0017] According to this embodiment, the determination unit can make the criteria for determining whether or not there is a deficiency in the character information different based on the combination of the first type and the second type, thereby improving the accuracy of extracting candidates for character information deficiencies.
[0018] In the embodiments described above, the first category may be a category relating to broadcasting.
[0019] According to this embodiment, it becomes possible to appropriately determine deficiencies in textual information contained in broadcast media data, according to the type of media.
[0020] In the embodiments described above, the first type may include at least one of the following: the data format of the media data, the output format of the media data, the genre of the program in which the media data is used, and the content of the media data.
[0021] According to this aspect, the accuracy of determining the deficiency of character information included in media data related to broadcasting is improved.
Effects of the Invention
[0022] According to the present invention, it is possible to appropriately determine the deficiency of character information included in a medium according to the type of the medium.
Brief Description of the Drawings
[0023] [Figure 1] It is a schematic diagram showing an example of the configuration of the information processing system 1 according to the present embodiment. [Figure 2] It is a diagram showing an example of the hardware configuration of the computer 100 according to the present embodiment. [Figure 3] It is a schematic diagram showing an example of the functional configuration of the server device 2 according to the present embodiment. [Figure 4] It is a schematic diagram showing an example of the details of the functional configuration of the extraction unit 30 according to the present embodiment. [Figure 5] It is a schematic diagram showing an example of the details of the functional configuration of the deficiency determination unit 40 and the output unit 50 according to the present embodiment. [Figure 6] It is a flowchart showing an example of the operation process of the server device 2 according to the present embodiment. [Figure 7] It is a schematic diagram showing an example of the screen output by the output unit 50 according to the present embodiment. [Figure 8] It is a schematic diagram showing another example of the screen output by the output unit 50 according to the present embodiment.
Modes for Carrying Out the Invention
[0024] A preferred embodiment of the present invention will be described with reference to the accompanying drawings. (In each figure, those denoted by the same reference numerals have the same or similar configurations.)
[0025] (1) Overall Configuration Figure 1 is a schematic diagram showing an example of the configuration of an information processing system 1 according to this embodiment. The information processing system 1 includes, for example, a server device 2 and a user terminal 3. The server device 2 is an example of an information processing device, and is an example of a system for determining deficiencies in character information contained in predetermined media data. Here, the media data is data from media such as manuscripts, flip charts, and captions. The media data will be described further later. The user terminal 3 is an example of an information processing device operated by a user. The server device 2 and the user terminal 3 are connected, for example, via a communication network N such as the Internet.
[0026] Server device 2, for example, acquires media data from user terminal 3, analyzes the media data to determine any deficiencies in the text information contained in the media data, and outputs the determination result. Server device 2 can target any media data for determination, but in particular, the media data may be related to broadcasting. Specifically, the media data may include data for on-screen text and flip charts in program production, and video data (complete package media, etc.) for broadcasting (regardless of whether it is in preparation for broadcasting, broadcasting, or secondary use).
[0027] (2) Hardware configuration Figure 2 shows an example of the hardware configuration of the computer 100 according to this embodiment. At least a part of the server device 2 and user terminal 3 according to this embodiment is each composed of, for example, one or more computers 100. Specifically, the computer 100 is envisioned to be a server device, a personal computer (PC), a PDA (Personal Digital Assistant), a mobile information terminal such as a smartphone or tablet terminal, a mobile phone, etc., but is not limited to these.
[0028] Computer 100 includes a processor 101, memory 102, storage 103, input / output interface (I / F) 104, and communication interface (Communication I / F) 105. Each hardware component of computer 100 is interconnected, for example, via bus B. Computer 100 realizes the functions and / or methods according to this embodiment through the cooperation of the processor 101, memory 102, storage 103, I / O I / F 104, and communication I / F 105.
[0029] The processor 101 executes functions and / or methods realized by code or instructions contained in a program stored in the storage 103. The processor 101 may include, for example, a central processing unit (CPU), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), a microprocessor, a processor core, a multiprocessor, an ASIC (Application-Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), etc., and may realize each of the embodiments by logic circuits (hardware) or dedicated circuits formed on an integrated circuit (IC (Integrated Circuit) chip, LSI (Large Scale Integration)), etc.
[0030] Memory 102 temporarily stores the program loaded from storage 103 and provides a workspace for the processor 101. Various data generated while the processor 101 is executing the program are also temporarily stored in memory 102. Memory 102 may include, for example, RAM (Random Access Memory) or ROM (Read Only Memory).
[0031] The storage 103 stores programs, for example. The storage 103 may include, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), or flash memory.
[0032] The input / output interface 104 has an input device for inputting various information to the computer 100, and an output device for outputting processing results processed by the computer 100. The input / output interface 104 may have the input device and output device integrated, or they may be separated into an input device and an output device.
[0033] The input device receives input from the operator and supplies the information related to that input to the processor 101. The input device may include, for example, hardware keys such as a touch panel, touch display, and keyboard, operation input devices such as a pointing device such as a mouse, image input devices such as a camera, and audio input devices such as a microphone.
[0034] The output device outputs the processing results processed by the processor 101. The output device may include, for example, a display device that can display images (including videos) according to display data (screen information) as a result of processing by the processor 101. The display device may include, for example, a touch panel, a touch display, a monitor (e.g., a liquid crystal display, an OLED (Organic Electroluminescence Display), etc.), a head-mounted display (HMD), projection mapping, a hologram, etc. The output device may also include, for example, an audio output device that can output audio according to audio data as a result of processing by the processor 101. The audio output device may include, for example, a speaker. The output device may also include, for example, a printing output device that outputs printed materials such as paper by printing.
[0035] The communication interface 105 transmits and receives various types of data between the information processing device and other information processing devices via the communication network N. This communication may be performed via wired or wireless (including short-range wireless communication), and any communication protocol may be used as long as communication between the devices is possible. The communication interface 105 has the function of performing communication with other information processing devices and equipment via the communication network N. The communication interface 105 transmits various types of data to other information processing devices according to instructions from the processor 101. The communication interface 105 also receives various types of data transmitted from other information processing devices and equipment and supplies them to the processor 101.
[0036] The program of this embodiment may be provided stored on a computer-readable storage medium. The storage medium is a “non-temporary, tangible medium” capable of storing the program. The program includes, for example, software programs and computer programs. The storage medium may, where appropriate, include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disk drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable storage medium, or two or more suitable combinations thereof. The storage medium may, where appropriate, be volatile, non-volatile, or a combination of volatile and non-volatile.
[0037] The program according to this embodiment may be provided to the computer 100 via any communication network N capable of transmitting the program. Furthermore, this embodiment can also be realized in the form of data signals embedded in a carrier wave, where the program is embodied through electronic transmission. The program according to this embodiment may be implemented using, for example, a scripting language such as ActionScript or JavaScript®, an object-oriented programming language such as Objective-C or Java®, or a markup language such as HTML5.
[0038] At least a portion of the processing in computer 100 may be implemented by cloud computing, which consists of one or more computers. At least a portion of the processing in computer 100 may be performed by other information processing devices. In this case, at least a portion of the processing of each functional unit implemented by processor 101 may be performed by other information processing devices.
[0039] (3) Functional configuration of server device 2 Referring to Figures 3 to 5, an example of the functional configuration of the server device 2 according to this embodiment will be described. Figure 3 is a schematic diagram showing an example of the functional configuration of the server device 2 according to this embodiment. Figure 4 is a schematic diagram showing an example of the detailed functional configuration of the extraction unit 30 according to this embodiment. Figure 5 is a schematic diagram showing an example of the detailed functional configuration of the defect determination unit 40 and output unit 50 according to this embodiment.
[0040] The server device 2 comprises a media data acquisition unit 10, a category determination unit 20, an extraction unit 30, a defect determination unit 40, and an output unit 50.
[0041] The media data acquisition unit 10 acquires media data, for example, by receiving media data as input to be used for determining deficiencies in character information. The media data acquisition unit 10 may acquire media data based on the input of media data to the input device provided by the server device 2, or it may acquire it by receiving media data from other devices outside the server device 2 (such as the user terminal 3 or a predetermined storage unit). The media data acquisition unit 10 may, for example, perform the necessary processing on the acquired media data and then supply it to other functional units such as the category determination unit 20 and the extraction unit 30.
[0042] Media data may include text information. "Including text information" is not limited to cases where text information is included as part of the media data, but may also include cases where text information can be generated based on processing such as image recognition processing of the media data. Media data may, for example, be related to broadcasting. The data format of the media data is not particularly limited, but may include, for example, images and text information. The output form of the media data is not particularly limited, but may include, for example, captions, flip charts, and scripts. The genre of programs in which the media data is used is not particularly limited, but may include, for example, sports, news, and information programs. The content of the media data is not particularly limited, but may include, for example, politics, economics, and baseball.
[0043] The category determination unit 20, for example, receives media data as input, determines the category of the media data (an example of a first type), and generates information indicating the category (category information) based on the determination. Here, the category of the media data (an example of a first type) may be any type of media data, and may include, for example, the data format of the media data (images, text information, etc.), the output format of the media data (captions, flips, and manuscripts, etc.), the genre of the program in which the media data is used (sports, news, and information programs, etc.), and the content of the media data (politics, economics, and baseball, etc.). The category determination unit 20 may also determine the category of the media data based on, for example, information about the environment of other applications that input the media data (for example, Google Workspace, Google Chat space information, or Slack channel information, etc.). The category determination unit 20 may also determine the category of the media data based on, for example, information entered by any user. The category determination unit 20 may also determine the category of the media data by, for example, analyzing the media data. The category determination unit 20 may, for example, supply category information to other functional units such as the extraction unit 30 and the defect determination unit 40.
[0044] The extraction unit 30 extracts, for example, at least one candidate (deficiency candidate) for character information defects based on media data. The extraction unit 30 outputs the deficiency candidate as a result of the extraction to the deficiency determination unit 40. The extraction unit 30 also supplies output data as a result of the extraction to the output unit 50. The extraction unit 30 may include, for example, a character information extraction unit 31, a character information processing unit 32, a typo DB 33, a typo DB determination unit 34, a typo DB deficiency candidate extraction unit 35, a system prompt list 36, a prompt determination unit 37, an LLM (Large Language Model) 38, and a language model deficiency candidate extraction unit 39.
[0045] The character information extraction unit 31 accepts media data such as images and PDFs as input, and generates character information from the images and PDFs using image recognition processing, OCR, etc. The character information extraction unit 31 may, for example, include position information indicating the position of characters in the media data in the character information it generates. The character information extraction unit 31 may, for example, output the generated character information to the character information processing unit 32 and the output unit 50, etc.
[0046] The character information processing unit 32 receives character information as input from, for example, the media data acquisition unit 10 and the character information extraction unit 31, and performs preprocessing on the character information. The content of the preprocessing is not particularly limited, but may include, for example, character division, removal of symbols, predetermined screening processes, etc. The character information processing unit 32 may output the preprocessed character information to, for example, the misspelling database defect candidate extraction unit 35 and the language model defect candidate extraction unit 39, etc.
[0047] The typographical error database 33 may be, for example, one database or a collection of multiple databases containing information about typographical errors. Here, "typographical error" may include not only simple character mistakes but also arbitrary notational defects such as "misspellings" and "omissions." "Mispellings" may include, for example, misuse or conversion errors of kanji or kana. "Omissions" may include, for example, the absence of words or characters that should be there. The typographical error database 33 may include information about typographical errors specific to broadcasting, and in addition to typographical errors, it may also include information about homonyms and variant characters as a warning. The typographical error database 33 may register information such as words that have been misspelled in the past, typographical errors that have recurred, words that were prevented at the manuscript stage, and words that are not necessarily errors but whose expression is prone to fluctuation. The typographical error database 33 may also include information to convey the importance to the user (typographical errors, warnings, cautions, etc.). The typographical error database 33 may also include additional information that is easy to handle with the above category information. The operation of the typographical error database 33 may be carried out under the management and cooperation of at least one department within the broadcasting station. Furthermore, the typographical error database 33 may be linked with other relevant data within the broadcasting station group. The typographical error database 33 may be updated at any time (including the addition of new typographical error information and the correction of existing typographical error information).
[0048] The typo DB determination unit 34, for example, receives category information as input from the category determination unit 20, and determines a reference typo DB (reference typo DB) from the typo DB 33 based on the category information. If the typo DB 33 consists of multiple DBs, the typo DB determination unit 34 may determine at least one DB from the multiple DBs as the reference typo DB. Alternatively, the typo DB determination unit 34 may filter the information contained in the typo DB 33 and determine the information extracted by the filtering as the reference typo DB. Note that one or more DBs contained in the typo DB 33, or the information contained in the typo DB 33, may have category information associated with them in advance, and the typo DB determination unit 34 may determine the reference typo DB by comparing the category information contained in the typo DB 33 with the category information received from the category determination unit 20. Alternatively, the typographical error database determination unit 34 may determine the reference typographical error database by, for example, comparing the category information determined by analyzing the information contained in the typographical error database 33 with the category information received from the category determination unit 20. The typographical error database determination unit 34 may, for example, supply the information of the reference typographical error database to the typographical error database defect candidate extraction unit 35 and / or the language model defect candidate extraction unit 39, etc.
[0049] The typographical error database defect candidate extraction unit 35 is an example of a first extraction unit, and extracts, for example, candidates (first defect candidates) for defects in the character information of media data. The typographical error database defect candidate extraction unit 35 generates first defect candidates based on pre-processed character information obtained from the character information processing unit 32, by referring to the information of the reference typographical error database supplied from the typographical error database determination unit 34. The typographical error database defect candidate extraction unit 35 may evaluate the possibility that each character unit (word, phrase, etc.) in the pre-processed string is a typographical error based on similarity to known typographical error patterns registered in the reference typographical error database, partial matching, or pattern matching such as regular expressions. The typographical error database defect candidate extraction unit 35 may then extract character units whose similarity exceeds a predetermined threshold as first defect candidates. Furthermore, the typographical error database defect candidate extraction unit 35 may add replacement candidate information and correction candidate information included in the reference typographical error database to the extracted first defect candidates.
[0050] The system prompt list DB36 stores, for example, multiple system prompts. A "system prompt" here refers to a guide indicating instructions or output strategies given to a large-scale language model (LLM) when extracting candidates for deficiencies in character information, and is used to improve processing consistency and accuracy. The system prompts recorded in the system prompt list DB36 may be associated with category information (types of media data).
[0051] The prompt determination unit 37, for example, receives category information as input and, based on the category information, determines a prompt from among multiple system prompts stored in the system prompt list DB 36 to be used for processing by the language model defect candidate extraction unit 39.
[0052] LLM38 is an example of a language model, specifically a large-scale language model (LLM) for extracting candidates for deficiencies in character information. LLM38 is responsible for generating candidates for deficiencies in character information based on a given prompt. An LLM may be a deep learning model with hundreds of millions of parameters and trained on hundreds of gigabytes or more of natural language data. An example of an LLM is gpt-4o. Other examples of language models include language models (not limited to large-scale) and machine learning models. Furthermore, if the media data is image data, a Vision Language Model (VLM) can be used, which takes image data as input and generates candidates for image data deficiencies based on a given prompt. A VLM takes both images and text as input and generates text and images based on their relationships. Thus, it is possible to select an appropriate language model or visual language model depending on the application and the type of media data.
[0053] The language model deficiency candidate extraction unit 39 is an example of a second extraction unit, and extracts candidates (second deficiency candidates) for deficiencies in the character information of media data. The language model deficiency candidate extraction unit 39 extracts second deficiency candidates for deficiencies in character information by, for example, providing the LLM 38 with a prompt containing character information supplied by the character information processing unit 32. The prompt may be a system prompt determined by the prompt determination unit 37 from among the system prompts included in the system prompt list DB 36. The language model deficiency candidate extraction unit 39 may also generate second deficiency candidates using the Retrieval-Augmented Generation (RAG) function after obtaining information from the reference typographical error DB determination unit 34. Here, Retrieval-Augmented Generation is a technique that supplements text generation with information from private or proprietary data sources or the results of web searches. Retrieval-Augmented Generation can be implemented by combining a search model configured to search a large dataset or knowledge base and an LLM that receives the retrieved data and generates an appropriate text response. The language model deficiency candidate extraction unit 39 calculates, for example, the semantic similarity between the output generated by LLM38 and the pre-processed character information, and evaluates the possibility of deficiencies for each string unit based on that similarity. Alternatively, the system may be configured to complementarily determine the confidence level of the second deficiency candidate using the similarity between known misspelling patterns and correction candidates included in the reference misspelling database. Furthermore, the second deficiency candidate can be extracted by quantitatively evaluating the semantic consistency between the instruction sentence included in the prompt and the target string using similarity.
[0054] The second set of deficiency candidates may include candidates for deficiencies in textual information and the type of deficiency associated with those candidates (an example of the second type). The type can be set arbitrarily, but may include, for example, nominalization, incomplete sentence structure (expressions using inversion, ellipsis, abrupt endings, etc.), abbreviated captions, unique notation rules of broadcasting stations, misspellings, omissions, inconsistencies in notation, prohibited broadcast terms, misuse of homophones, names of people, titles / positions, place names, country names, missing particles, comments on readings, full-width / half-width characters, periods / commas, symbols, etc. A concrete example of a second set of deficiency candidates would be, for example, in the case of textual information such as "Tokyo Heavy Rain," the deficiency content "The sentence is incomplete," the appropriate expression "It is raining heavily in Tokyo...", and the type "Missing particle." Furthermore, other specific examples of potential deficiencies in the second category may include, for example, in the case of textual information such as "West Virginia," the deficiency content being "incorrect wording," the correct expression "West Virginia," and the type being "notation guidelines."
[0055] The server device 2 may include multiple language model deficiency candidate extraction units 39 and / or LLM 38. In particular, at least one language model deficiency candidate extraction unit 39 and / or LLM 38 may have a function to perform fact-checking by performing a web search or the like. The extraction unit 30 may generate a final second deficiency candidate by comparing multiple second deficiency candidates from multiple language model deficiency candidate extraction units 39 and / or LLM 38.
[0056] The deficiency determination unit 40 is an example of a determination unit that, for example, determines deficiencies in the character information contained in the media data based on category information obtained from the category determination unit 20 and at least one deficiency candidate (first deficiency candidate and / or second deficiency candidate) obtained from the extraction unit 30. The deficiency determination unit 40 generates output data as a result of the determination and supplies it to the output unit 50. The deficiency determination unit 40 comprises, for example, a deficiency determination processing unit 41, an LLM 42, and a determination information integration unit 43.
[0057] The deficiency determination processing unit 41 accepts, for example, deficiency candidates (first deficiency candidate and / or second deficiency candidate) as input and generates a deficiency candidate list based on the category information and the deficiency candidates (first deficiency candidate and / or second deficiency candidate). The deficiency candidate list may be, for example, a list indicating what kind of deficiencies exist in the character information of the media data based on the deficiency candidates (first deficiency candidate and / or second deficiency candidate). The deficiency determination processing unit 41 may, for example, perform a predetermined scoring on the deficiency candidates (first deficiency candidate and / or second deficiency candidate) and output the result as the final deficiency candidate list. The deficiency determination processing unit 41 may, for example, make different determinations on whether or not there are deficiencies in the character information based on a combination of category information and type information.
[0058] Furthermore, the defect detection processing unit 41 can switch the list of defect candidates generated according to the category information. For example, in general text such as manuscripts, types of defects such as "ending sentences with nouns" and "incomplete sentences" may be judged as defects. However, in the media industry, such as broadcasting, it is common to intentionally use "ending sentences with nouns" and "incomplete sentences" in media data such as flip charts and captions in order to prioritize readability and efficiency of information transmission. Taking these media-specific expression styles into consideration, the defect detection processing unit 41 can switch its processing so that these expressions are not judged as defects when the category information is flip charts or captions.
[0059] Furthermore, if the category information is news programs and the type of deficiency is "person's name" or "title / position," it is necessary to obtain the latest information. Therefore, the search extension generation function may be used to check the content of the deficiency candidates and generate a list of deficiency candidates. Alternatively, users may be prompted to "check as the deficiency concerns a person's name or title."
[0060] Furthermore, regarding media-related matters such as broadcasting, when rephrasing trademarks, it may be necessary to replace specific trademarks with common nouns. Also, regarding terminology unsuitable for broadcasting, care must be taken in selecting terms to avoid inappropriate expressions and discriminatory language. Regarding differences in notation depending on the program genre, the appropriate notation style differs depending on the program genre, such as news, variety, and documentary. Examples of compliance with notation rules of various industry associations include "garbage" and "trash," and "West Virginia" and "West Virginia." The deficiency judgment processing unit 41 can take these broadcast-specific circumstances into consideration and appropriately adjust the generation of deficiency candidate lists and judgment criteria based on category information.
[0061] Furthermore, the defect detection processing unit 41 has a metadata processing function, which allows it to process metadata that is not directly related to the content, such as creation IDs, creator information, and technical instructions such as size and placement, that are attached to flips and captions, and not to be judged as defects. This enables efficient defect detection that focuses on the essential content of the media data. For example, production information such as "#ID:T20250519-001" and "Size: Standard, Position: Bottom Right" included in caption data is management information, not content, and is therefore automatically excluded from defect detection.
[0062] Furthermore, the deficiency detection processing unit 41 has a function to flexibly control which types of content are not judged as deficiencies based on external settings. This makes it possible to customize the judgment criteria according to the unique style guides and notation rules of each broadcasting station or production company. For example, if a particular program series intentionally uses colloquial language, the system administrator can set it to exclude the "colloquial language" type from the deficiency detection criteria. Similarly, in content that is highly seasonal or topical, it may be necessary to temporarily allow certain expressions, so it is possible to temporarily disable the deficiency detection of certain types of content. In this way, the deficiency detection processing unit 41 can flexibly respond to the diverse needs of the broadcasting production field.
[0063] LLM42 is an LLM for generating an output statement that integrates a list of potential defects, for example, in response to a prompt based on a list of potential defects. The output statement is output data and may include, for example, a summary of the potential defects and a description of the defects (first and / or second potential defects).
[0064] The judgment information integration unit 43, for example, receives a list of deficiency candidates as input from the deficiency determination processing unit 41 and inputs it to the LLM 42 to generate an output statement that integrates the list of deficiency candidates. The judgment information integration unit 43 then supplies the generated output statement (output data) to the output unit 50.
[0065] The output unit 50 outputs the judgment result from the defect determination unit 40, for example, based on the output data obtained from the defect determination unit 40. The output unit 50 may also perform output processing based on the output data obtained from the extraction unit 30, for example. The output unit 50 includes, for example, a result image generation unit 51 and an output processing unit 52.
[0066] The result image generation unit 51 may, for example, acquire character information from the extraction unit 30. This character information may include location information (for example, location information indicating the position of characters in the media data). The result image generation unit 51 may, for example, acquire a list of defect candidates from the defect determination processing unit 41. The result image generation unit 51 may, for example, generate a result image based on the acquired character information and / or the list of defect candidates. A result image is a visual output created based on the character information and location information contained in the media data and the list of defect candidates. The image may be generated with the defect highlighted or otherwise indicated so that the user can see at a glance where and what kind of defect there is in the character information.
[0067] The output processing unit 52 may, for example, output a judgment result based on the output text obtained from the judgment information integration unit 43 and / or the result image obtained from the result image generation unit 51. The output processing unit 52 may, for example, control an output device (display device and / or audio output device, etc.) provided by the server device 2 to output the judgment result (display output and / or audio output, etc.). Alternatively, the output unit 50 may, for example, transmit the judgment result to another device (user terminal 3, etc.) located outside the server device 2 as part of the judgment result output processing. The other device (user terminal 3, etc.) may control a predetermined output device (display device and / or audio output device, etc.) to output the judgment result received from the server device 2 (display output and / or audio output, etc.).
[0068] (4) Operation processing Figure 6 is a flowchart showing an example of the operation process of the server device 2 according to this embodiment. The process shown in this figure illustrates the sequence of steps from when the server device 2 detects a defect in the character information contained in the media data to when it outputs the judgment result.
[0069] First, in step S1, the server device 2 acquires media data. The media data may be, for example, data input from the user terminal 3 or data stored in a storage device or the like.
[0070] Next, in step S2, the category determination unit 20 determines the category of the media data. This category determination is performed based on, for example, the intended use and display format of the media data, and is used for subsequent processing.
[0071] In the following step S3, a process for extracting potential deficiencies is performed. Specifically, in step S3-1, the typographical error database deficiency candidate extraction unit 35 extracts a first deficiency candidate by referring to the typographical error database, and in step S3-2, the language model deficiency candidate extraction unit 39 extracts a second deficiency candidate using a large-scale language model (LLM). Note that the order of steps S3-1 and S3-2 does not matter; either one may be executed first, or both may be executed in parallel.
[0072] Subsequently, in step S4, the defect determination unit 40 determines whether there are defects in the character information based on these defect candidates. The extracted candidate information and category information are used for the determination, and the presence and content of defects are evaluated.
[0073] Finally, in step S5, the output unit 50 outputs the judgment result. The output is provided, for example, as text containing the details of the defects or as a result image highlighting the defective areas.
[0074] (5) Example output Figure 7 is a schematic diagram showing an example of a screen output by the output unit 50 according to this embodiment. Screen G1 is shown in the figure as an example of such a screen.
[0075] Screen G1 is a screen for displaying deficiencies in text information based on media data. Screen G1 may be displayed, for example, on user terminal 3. Screen G1 includes, for example, areas G11, G12, and G13.
[0076] Area G11 may include, for example, a summary of deficiencies in the text information. In the illustrated example, area G11 includes text indicating the judgment result, "Several typographical errors were found!" Area G11 also includes text that draws attention to deficiencies requiring particular attention, such as, "Please check the date notation, especially for 'December 31st' and '1st 26th'," and "Also, 'Sattomi' needs to be corrected to 'Summit,' and 'Hoiketsu' might be better changed to 'Drilling'." These items correspond to output statements generated by, for example, the deficiency judgment unit 40 and the judgment information integration unit 43.
[0077] The items in area G12 may include, for example, deficiencies corresponding to the first candidate deficiency (extracted by the Typo Database Deficiency Candidate Extraction Unit 35) extracted based on the typographical error database (Typographical Error DB 33). The first point in area G12 is that the characters "1st 26th" included in the media data are a notational error, pointing out the possibility of an error in the date notation. The second point in area G12 is that the characters "Sattomi" included in the media data are a notational error, pointing out an error in the katakana notation of the name of the G7 Summit, and suggesting a correction to "Summit". The third point in area G12 is that the characters "Horekaku" included in the media data are a misuse of kanji, and a correction to "Excavation" is suggested. The fourth point in area G12 is that the characters "Weekly Paper" included in the media data are a mix-up of words, and a message is provided prompting confirmation of whether a correction to "Weekly Magazine" is necessary.
[0078] The items in area G13 may include, for example, deficiencies corresponding to the second deficiency candidate extracted by the language model deficiency candidate extraction unit 39. The first point in area G13 concerns the phrase "house intrusion" included in the media data. It is suggested that the expression of the phrase is inappropriate, and that the meanings of both the misspelled "intrusion" and the correct spelling "intrusion" be corrected to "intrusion into a house." The second point in area G13 concerns the phrase "fields and rice paddies" included in the media data. It is suggested that the expression lacks generality and that it be replaced with "fields and rice paddies." The third point in area G13 concerns the phrase "base" included in the media data. It is suggested that the meaning of the phrase be incorrect and that it be corrected to "cemetery."
[0079] Figure 8 is a schematic diagram showing another example of a screen output by the output unit 50 according to this embodiment. Screen G2 is shown in the figure as an example of such a screen.
[0080] Screen G2 is a screen for displaying deficiencies in text information based on media data. Screen G2 may be displayed on, for example, user terminal 3. Screen G2 includes, for example, a result image, and the illustrated example is an example of determining media data for a news program script.
[0081] Area G21 contains textual information from the script for the news program. Each line of text is enclosed in a rectangular frame. In particular, the following sentences contain errors: "flooding due to rising river levels," "animals entering houses due to landslides," "don't go to check on the irrigation canals in the fields," and "this year due to a decrease in El Niño." Specifically, "flooding" should be "flooding," "intrusion" should be "intrusion," "fields" should be "fields," and "decrease in El Niño" should be "El Niño phenomenon." As shown in the diagram, these sentences in area G21 have thicker borders than other sentences to highlight the parts containing errors. Note that in screen G2, in addition to thickening the borders, errors may be highlighted in any manner, such as by changing the color of the text, underlining, or changing the background color of the erroneous parts.
[0082] As described above, according to this embodiment, a determination of character information defects is made based on a first type, which is the type of media data, and at least one candidate for defects in character information contained in the media data, and the determination result is output. Therefore, it becomes possible to appropriately determine character information defects contained in the media according to the type of media.
[0083] It should be noted that the present invention is not limited to the embodiments described above, and can be implemented in various other forms without departing from the spirit of the invention. For this reason, the above embodiments are merely illustrative in all respects and should not be interpreted restrictively. For example, the order of each processing step described above can be arbitrarily changed or executed in parallel, as long as there is no inconsistency in the processing content. [Explanation of symbols]
[0084] 1…Information processing system, 2…Server device, 3…User terminal, 10…Media data acquisition unit, 20…Category determination unit, 30…Extraction unit, 31…Character information extraction unit, 32…Character information processing unit, 33…Typo DB, 34…Typo DB determination unit, 35…Typo DB defect candidate extraction unit, 36…System prompt list, 37…Prompt determination unit, 38…LLM, 39…Language model defect candidate extraction unit, 40…Defect determination unit, 41…Defect determination processing unit, 42…LLM, 43…Determination information integration unit, 50…Output unit, 51…Result image generation unit, 52…Output processing unit, 100…Computer, 101…Processor, 102…Memory, 103…Storage, 104…Input / Output Interface (Input / Output I / F), 105…Communication Interface (Communication I / F)
Claims
1. A media data acquisition unit that acquires media data, A first type determination unit that determines a first type which is the type of the media data based on the media data, An extraction unit that extracts at least one candidate for a deficiency in the character information contained in the aforementioned media data, A determination unit that determines whether the character information is incomplete based on the first type and the at least one candidate, An information processing apparatus comprising an output unit that outputs the determination result of the determination unit, The determination unit is an information processing device that makes the criteria for determining whether or not the character information is defective different based on the combination of the first type and the second type, which is the type of defect in the character information.
2. The extraction unit is A first extraction unit extracts a first candidate from the at least one candidate by referring to a predetermined typographical error database, A second extraction unit extracts a second candidate from the at least one candidate by inputting a predetermined prompt into a predetermined language model, The information processing apparatus according to claim 1, comprising at least one of the following.
3. The information processing apparatus according to claim 2, wherein the extraction unit further comprises a typographical error database determination unit that determines a predetermined typographical error database from a plurality of typographical error databases based on the first type.
4. The information processing apparatus according to claim 2, wherein the extraction unit further comprises a prompt determination unit that determines a predetermined prompt from a plurality of prompts that can be input to a predetermined language model based on the first type.
5. The information processing apparatus according to claim 2, wherein the second extraction unit extracts the second candidate associated with the second type.
6. The information processing apparatus according to claim 1, wherein the first type is a type related to broadcasting.
7. The information processing apparatus according to claim 6, wherein the first type includes at least one of the data format of the media data, the output form of the media data, the genre of the program in which the media data is used, and the content of the media data.
8. One or more computers, Acquiring media data, Based on the media data, determine the first type, which is the type of the media data. Extracting at least one candidate for a deficiency in the textual information contained in the aforementioned media data, Based on the first type and the at least one candidate, a determination is made as to whether the character information is incomplete. An information processing method that performs the following: outputting the result of the determination, The method for performing the aforementioned determination involves differentiating the criteria for determining whether or not the character information is defective based on the combination of the first type and the second type, which is the type of defect in the character information.
9. One or more computers, A media data acquisition unit that acquires media data, A first type determination unit that determines a first type which is the type of the media data based on the media data, An extraction unit that extracts at least one candidate for a deficiency in the character information contained in the aforementioned media data, A determination unit that determines whether the character information is incomplete based on the first type and the at least one candidate, A program for causing an output unit to function as an output unit that outputs the determination result of the determination unit, The determination unit is a program that makes the criteria for determining whether or not the character information is defective differ based on the combination of the first type and the second type, which is the type of defect in the character information.