Information processing system, information processing method and program

The information processing system addresses the challenge of unclear outputs by transforming semantic vectors into meaningful words and phrases, enhancing user comprehension of the data.

JP7757666B2Active Publication Date: 2025-10-22KONICA MINOLTA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021146707
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-09
Publication Date
2025-10-22
Estimated Expiration
2041-09-09

AI Technical Summary

Technical Problem

Existing information processing systems struggle to determine correspondences and dependencies between selected terms, leading to outputs that are often mysterious or incomprehensible, making it difficult for users to interpret the data.

Method used

An information processing system that converts input data into semantic vectors, performs calculations to transform these vectors into different directions, and then converts them back into text, extracting meaningful words and phrases to generate output data.

Benefits of technology

The system outputs data that is easier for users to interpret by eliminating unclear dependencies and redundancies, allowing for clearer understanding of the content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007757666000001
    Figure 0007757666000001
  • Figure 0007757666000002
    Figure 0007757666000002
  • Figure 0007757666000003
    Figure 0007757666000003
Patent Text Reader

Abstract

To provide an information processing system, an information processing method, and a program, which are capable of outputting data that is easier for a user to interpret.SOLUTION: An information processing system includes: input means for receiving input data having a plurality of elements; acquisition means for acquiring one first semantic vector having a direction and indicating a content of the input data by the direction; calculation means for performing an arithmetic operation to convert the acquired first semantic vector into a second semantic vector having a direction different from that of the first semantic vector; and generation means for generating output data acquired by extracting a plurality of word phrases included in a content of the second semantic vector.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] In recent years, advances in natural language processing have led to the development of information processing systems that respond in natural language to various input data. Patent Document 1 discloses a technology that supports user ideas by extracting word pairs from input sentences, synthesizing and outputting sentences containing the obtained word pairs after appropriate transformation and conversion processing depending on the case. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-86815 Summary of the Invention [Problem to be solved by the invention]

[0004] However, at present, it is difficult to mechanically determine the correspondences and dependencies between a wide range of selected terms regardless of syntax, etc., and the output as a whole is often mysterious or incomprehensible, making it impossible for users to interpret.

[0005] An object of the present invention is to provide an information processing system, an information processing method, and a program that can output data that is easier for the user to interpret. [Means for solving the problem]

[0006] In order to achieve the above object, the invention described in claim 1 is: Multiple No. 1 input means for accepting input data having elements; an acquisition means for acquiring a first semantic vector having a direction and indicating the content of the input data by the direction; The acquired first semantic vector is a vector having a different direction from the first semantic vector. Single A calculation means for performing a calculation to convert the data into a second semantic vector; The aforementioned Single Second Semantic Vector The text is converted back into a text having a plurality of second elements, and the text is divided into a plurality of words, and a meaningful expression is obtained based on the words. Multiple words As the second elements of the plurality generating means for generating the extracted output data; The information processing system is characterized by comprising:

[0007] The invention described in claim 2 is the information processing system described in claim 1, the generating means extracts the plurality of words and phrases included in a series of texts, the content of which is the second semantic vector; The run of text is an incomplete sentence It is characterized by:

[0008] The invention of claim 3 provides the information processing system of claim 1 or 2, The plurality of phrases are characterized by being words or phrases.

[0009] The invention of claim 4 provides an information processing system according to any one of claims 1 to 3, The input data includes at least one of image data, text data, and audio data.

[0010] The invention of claim 5 is the information processing system of claim 4, the input data includes text data; The text data is sentence data. It is characterized by:

[0011] The invention of claim 6 provides an information processing system according to any one of claims 1 to 5, the input data includes non-text data; The non-text data is No. 1 A conversion means is provided to convert the data into text data containing elements. It is characterized by:

[0012] The invention of claim 7 provides an information processing system according to any one of claims 4 to 6, the input data includes image data; The image data includes at least one of photograph data, picture data, graphic data, and character data. It is characterized by:

[0013] The invention of claim 8 is an information processing system according to any one of claims 4 to 6, the input data includes audio data; The audio data includes at least one of speech and conversation data. It is characterized by:

[0014] The invention described in claim 9 is as follows: An information processing method by a control unit of a computer, Multiple No. 1 an input step of accepting input data having elements; an acquisition step of acquiring one first semantic vector having a direction and indicating the content of the input data by the direction; The acquired first semantic vector is a vector having a different direction from the first semantic vector. Single a calculation step of converting the vector into a second semantic vector; The aforementioned Single Second Semantic Vector The text is converted back into a text having a plurality of second elements, and the text is divided into a plurality of words, and a meaningful expression is obtained based on the words. Multiple words As the plurality of second elements a generating step for extracting and generating output data; characterized in that it includes do.

[0015] The invention described in claim 10 is as follows: Computer Multiple No. 1 input means for accepting input data having elements; an acquisition means for acquiring a first semantic vector having a direction and indicating the content of the input data by the direction; The acquired first semantic vector is a vector having a different direction from the first semantic vector. Single A calculation means for performing a calculation to convert the vector into a second semantic vector; The aforementioned Single Second Semantic Vector The text is converted back into a text having a plurality of second elements, and the text is divided into a plurality of words, and a meaningful expression is obtained based on the words. Multiple words As the plurality of second elements generating means for extracting and generating output data; The program is characterized by functioning as follows. [Effects of the Invention]

[0016] According to the present invention, it is possible to output data that is easier for the user to interpret. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a configuration diagram showing an information processing system according to an embodiment of the present invention; [Figure 2] FIG. 10 illustrates an example of processing a string returned from a semantic vector. [Figure 3] 10 is a flowchart showing a control procedure of a conversion output control process. [Figure 4] 10 is a flowchart showing a control procedure of an idea support control process. DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a configuration diagram showing an information processing system 1 of this embodiment. The information processing system 1 includes an information processing device 10, a database device 20, and a terminal device 30. The information processing system 1 is not particularly limited, but for example, the information processing device 10 performs an idea generation support process based on an input operation from the terminal device 30, and the process result is output to the terminal device 30 so that the result can be obtained by the terminal device 30.

[0019] The information processing device 10 is a typical server computer that includes a control unit 11 (acquisition means, calculation means, generation means, conversion means), a storage unit 12, a communication unit 13 (input means), etc. The control unit 11 has a hardware processor such as a CPU (Central Processing Unit) and controls the overall operation of the information processing device 10. The control unit 11 may have multiple CPUs, which may perform common processing in parallel or may each perform calculation processing independently for each purpose.

[0020] The storage unit 12 has a RAM (Random Access Memory) and a non-volatile memory. The RAM provides a working memory space for the CPU and stores temporary data. The non-volatile memory stores and holds the program 121, setting data, etc. The non-volatile memory is, for example, a flash memory, but may also include an HDD (Hard Disk Drive), etc. The non-volatile memory may also include a network drive, cloud storage, etc.

[0021] The program 121 includes a digitization processing unit 1211 and a text conversion unit 1212. The digitization processing unit 1211 converts text data into a multidimensional vector (semantic vector) or converts a multidimensional vector into text data. The text conversion unit 1212 recognizes image data and audio data and converts them into sentence text data.

[0022] The communication unit 13 includes a network card that controls communication with external devices, etc. The communication unit 13 may be capable of controlling communication with external devices based on, for example, a LAN (Local Area Network) standard.

[0023] The database device 20 is a database that stores correspondence data for the digitization processing unit 1211 to convert between text and semantic vectors, and includes a storage unit 21. That is, the storage unit 21 stores term data 211 that associates text (phrases and sentences) with semantic vectors.

[0024] The database device 20 can be accessed from the information processing device 10 via the communication unit 13 on the network.

[0025] The terminal device 30 requests the information processing device 10 to perform processing, and also receives and acquires processing results from the information processing device 10. The terminal device 30 includes a control unit 31, a storage unit 32, a display unit 33 (output means), an operation reception unit 34, a communication unit 35, etc.

[0026] The control unit 31 has a CPU and controls the overall operation of the terminal device 30 . The storage unit 32 includes a RAM and a non-volatile memory. The RAM provides a working memory space for the CPU and stores temporary data. The non-volatile memory stores control programs, setting data, and the like.

[0027] The display unit 33 has a display screen and performs display operations on the display screen under the control of the control unit 31. The display screen is, for example, a liquid crystal display (LCD) and can display characters, figures, and images within the range that can be displayed with the resolution of the LCD. The display unit 33 may also have an LED lamp for indicating the presence or absence of power supply, the status of access to the memory unit 32, the charging status, etc.

[0028] The operation reception unit 34 receives an input operation from outside, generates an operation signal, and outputs the operation signal to the control unit 31. The operation reception unit 34 includes, for example, a keyboard and a pointing device (such as a mouse). Additionally or alternatively, the operation reception unit 34 may have a touch panel positioned so as to overlap the display screen, and may detect a touch operation and generate an operation signal.

[0029] The communication unit 35 controls communication with external devices and is capable of communicating with the information processing device 10 via a LAN.

[0030] Next, the idea generation support operation in the information processing system 1 will be described. In the information processing system 1, a trained model for distributed representation (multidimensional vectorization) of documents is stored in advance in the storage unit 21 of the database device 20, and input data input (accepted) from the terminal device 30 via the communication unit 13 is converted into a multidimensional vector using the trained model. This multidimensional vector is a semantic vector having a direction according to meaning, and can be calculated. For example, adding two semantic vectors results in a vector direction having the two meanings. The number of dimensions may be determined appropriately, but is, for example, approximately 50-200. By performing calculations on this semantic vector, it is converted into a direction different from the original semantic vector, and the converted vector is then converted back into text and output, thereby obtaining output content different from the input content. Depending on the type of calculation, content based to some extent on the input content or content different from the input content can be obtained.

[0031] Doc2vec, an extension of Word2Vec, is known as a technology for semantic vectorizing (obtaining a first semantic vector) the content of a sentence (input data, text data) having multiple elements (phrases) each with their own individual meaning. Doc2Vec is a technology that breaks down the content of an input document into words and phrases, and uses an unsupervised trained model that has been trained to create a distributed representation while taking into account the frequency of appearance, combinations, and order of appearance of each word (although some algorithms do not take this into account). The breakdown into words and phrases can be performed using a conventionally well-known technology, such as morphological analysis.

[0032] The acquired input data does not necessarily have to be text data. If the input data is non-text data, such as image data or audio data, the text conversion unit 1212 may perform a processing operation to recognize the content of the data and convert it into text data, after which the same processing as for text data may be performed.

[0033] The voice data may be processed to recognize words uttered in conversations or speeches and sequentially convert them into text. In this case, rather than simply converting words into text, the text data may also include information such as speaking habits, voice quality, volume, and speaking rate (which may be relative changes). This additional information may be made identifiable by a specific format, such as parentheses. Voice quality may be determined, for example, based on voice frequency analysis. Since volume also depends on the recording conditions, it may be determined based on relative differences within a conversation and / or between multiple speakers, rather than absolute values. Speaking rate may be determined based on the average length of each sound.

[0034] With regard to image data, if the content of the image data is text (character shape data), each character shape can be recognized and converted into text using well-known OCR (Optical Character Reader) technology. In this case, the text data may also include design information such as font size, font type, and character color, as well as information such as vertical or horizontal writing and variations in handwriting. Furthermore, if the content of the image data is a figure, painting, landscape, or a photograph of these, well-known image recognition processing (e.g., using a convolutional neural network) is performed to recognize the content of the image and convert it into corresponding text. In this case, the text data may also include information such as the recognized content, its position and size (proportion of the image), overlap, and positional relationship.

[0035] The input data may also be a combination of multiple types of data, such as text data accompanied by drawings. In this case, the image portion may be recognized and converted into text, and the resulting text may be inserted into the text corresponding to the image. The input data may also be a combination of multiple independent input contents, in which case the text of two pieces of input data may simply be arranged in order.

[0036] The obtained semantic vector is then subjected to a calculation to transform it in an appropriate different direction to obtain a transformed semantic vector (second semantic vector). Since the second semantic vector obtained for idea generation support is meaningless if it is too similar to the first semantic vector, the transformation direction may be determined so that a predetermined difference is obtained, for example, using cosine similarity or Euclidean distance. It may also be possible to significantly change some of the components of the multidimensional vector, or conversely, fix the components with large values ​​and change the values ​​of the other components. Alternatively, a reference multidimensional vector may be added or subtracted from the obtained first semantic vector.

[0037] Semantic vectors, which are distributed representations of sentences in this way, are compressed into a number of dimensions set according to the content category, etc., and are therefore suitable as indices for quantitatively evaluating the similarity of sentences and extracting and outputting the resulting similar documents. However, it is difficult to convert these semantic vectors back into sentences, making it difficult to obtain sentences that can be properly understood by humans as natural language expressions.

[0038] In the information processing system 1, the semantic vector obtained is converted back (inversely converted) into data of a character string (a string of text), and then multiple meaningful words that appear in this character string, which is usually an incomplete sentence, are extracted and output data is generated in which the words are separated and listed. By outputting this output data, unnatural sentences, meaningless dependencies, meaningless character strings, duplicated characters, etc. are excluded from the output.

[0039] FIG. 2 is a diagram showing an example of processing a string returned from a semantic vector. Even if the semantic string output as shown in Figure 2(a) does not make sense as a sentence, as shown in Figure 2(b), it is first divided into words using morphological analysis, and then independent words are selected by excluding adjuncts and redundant phrases, etc. From these, nouns and nouns that form the stems of sa-hen verbs may be further extracted depending on the situation, as shown in Figure 2(c).

[0040] In this way, by extracting only meaningful phrases (words) and outputting them in a row, it becomes easier to understand what combination of phrases the semantic vector represents. Note that if there is a part in the string that makes sense as a phrase, the phrase may be output as a whole without being separated.

[0041] The extracted data to be output is sent by the control unit 11 to the terminal device 30 via the communication unit 13, and is output by being displayed on the display unit 33 under the control of the control unit 31, although this is not particularly limited.

[0042] 3 is a flowchart showing a control procedure by the control unit 11 of the conversion output control process executed by the information processing device 10 of this embodiment. This conversion output control process is started when input data is received from the terminal device 30 via the communication unit 13 or the like.

[0043] When the conversion output control process starts, the control unit 11 (CPU) accepts and acquires received input data (step S101; input means, input step). The control unit 11 determines whether the acquired input data is text data (step S102). If it is determined that the input data is not text data ("NO" in step S102; including cases where some data is not text data), the control unit 11 converts the non-text portion of the input data into text document data using the text conversion unit 1212 (step S103). Then, the control unit proceeds to step S104. If it is determined that the input data is text data ("YES" in step S102), the control unit proceeds to step S104.

[0044] When the process proceeds to step S104, the control unit 11 converts the text document data into a semantic vector with a predetermined number of dimensions (step S104; acquisition means, acquisition step). The conversion may be performed, for example, using doc2vec as described above. The control unit 11 then performs a predetermined calculation process on the obtained semantic vector (step S105; calculation means, calculation step). The calculation process may involve, for example, randomly acquiring one of the directions in which the cosine similarity indicates a predetermined value while maintaining some components of the semantic vector.

[0045] The control unit 11 reverse-converts the semantic vector obtained by the calculation into a series of text (step S106). As described above, it is generally difficult to obtain text obtained by reverse conversion that accurately expresses natural language, and the meaning may be unclear or unnatural dependencies may occur. The control unit 11 breaks the obtained text into words and phrases using morphological analysis or the like, and generates output data in which words and phrases that satisfy conditions are extracted (step S107; generation means, generation step). As described above, the words and phrases that satisfy the conditions may, for example, simply be extracted as independent words, or may be limited to nouns, etc. Furthermore, adjectives, verbs, etc. may be converted to their basic forms as appropriate and extracted.

[0046] Control unit 11 returns the list of extracted words and phrases to the sender of the input data via communication unit 13 (step S108), and control unit 11 then ends the conversion output control process.

[0047] FIG. 4 is a flowchart showing a control procedure by the control unit 31 of the idea support control process executed by the terminal device 30.

[0048] When the idea generation support control process is started, the control unit 31 sets input data that will be the basis for ideas (step S301). The setting is performed, for example, by accepting a command to specify a document data file, an image data file, or an audio data file, or to select a part of a document, an image, or the like displayed on the display unit 33, based on a user's input operation to the operation acceptance unit 34. Alternatively, the control unit 31 may set a sentence entered in a text entry field as input data, or may record what the user says via a microphone or the like connected to an audio input terminal (not shown) and use the recorded data as input data.

[0049] The control unit 31 transmits the set input data to the information processing device 10 via the communication unit 35 (step S302). The control unit 31 waits for a reply from the information processing device 10, and receives reply data from the information processing device 10, i.e., the list data of words and phrases output in step S108 of the conversion output control process (step S303).

[0050] When the control unit 31 receives the list data of the words and phrases, it causes the display unit 33 to perform an output operation by displaying a list of the words and phrases included in the list data (output means, output step). Then, the control unit 31 ends the idea generation support control process.

[0051] Note that meaningful words and phrases do not have to be simply extracted from a series of text obtained by inverse transformation. For example, if sentence data with semantic vectors oriented similar to the obtained semantic vector is stored, the words and phrases in the sentence may be compared with the extracted words and phrases, and words and phrases with overlapping meanings may be further selected and output. The directional closeness between the semantic vectors is not particularly limited, and may be determined by, for example, cosine similarity or Euclidean distance. The words and phrases included in the sentence data may be listed in advance using, for example, a bag of words (BoW). Among the extracted words and phrases, words and phrases that are frequently used in BoW or characteristic words may be further selected and output.

[0052] As described above, the information processing system 1 of this embodiment includes a communication unit 13 that receives input data having multiple elements (meaningful words, figures, etc.), a control unit 11, and a display unit 33. The control unit 11, as an acquisition unit, acquires a first semantic vector that has a direction and indicates the content of the input data based on the direction, as a calculation unit, performs a calculation to convert the acquired first semantic vector into a second semantic vector that has a direction different from that of the first semantic vector, and as a generation unit, extracts multiple words and phrases included in the content of the second semantic vector. When attempting to convert a semantic vector (second semantic vector) that relates multiple elements into a series of text, such as a sentence, it is difficult to obtain a sentence with natural and appropriate dependency and grammatical structure due to the characteristics of the semantic vector. Therefore, by displaying a list of phrases in the text converted from the semantic vector (here, "phrase" refers primarily to a minimum element with meaning, such as an independent word, or a minimum length sequence of multiple elements organized into a fixed expression such as a phrase; it does not refer to a general sentence or phrase in which elements (such as words) with different meanings are arbitrarily strung together), the user can more clearly grasp only the main points of the content, making it easier for the user to interpret the output content. Furthermore, since the output does not need to be complete, meaningful output can be obtained for some applications (such as idea generation support) even if the learning capacity of semantic vector conversion such as doc2vec is somewhat insufficient.

[0053] Furthermore, the control unit 11, as a generating means, extracts a plurality of words and phrases contained in the series of texts, treating the content of the second semantic vector as a series of texts. At this time, the series of texts is an incomplete sentence. As described above, even if an attempt is made to obtain a sentence from a semantic vector, an appropriate expression as a natural language expression cannot usually be obtained. Therefore, by extracting words, phrases, and other such phrases from the incomplete sentences and simply displaying them while eliminating unclear dependency relationships and redundancies, the information is easily organized, making it easier for the user to understand the output content.

[0054] Furthermore, the multiple phrases may be words or phrases. In this way, the information processing system 1 eliminates connections between non-standard words and simply outputs data in the smallest units, making it easier for the user to understand the output content.

[0055] The input data includes at least one of image data, text data, and audio data. That is, the input data does not have to be text data from the beginning, as long as it can be converted into a semantic vector as appropriate, so a wide range of input data can be handled.

[0056] The input data may also include text data, which may be sentence data. In this way, even if an entire sentence is input, a semantic vector quantified according to the content type of the sentence can be obtained, and the semantic vector can be appropriately converted to obtain output data containing words and phrases corresponding to the semantic vector. This eliminates the need for the user to sort through the contents of the input data one by one, and allows the user to easily obtain the gist of the content obtained based on sentences that are of interest or interest to the user.

[0057] Furthermore, the input data may include non-text data, and the control unit 11 may function as a conversion means to convert the non-text data into text data including the multiple elements. In other words, when non-text data is input, by converting it into text data and then obtaining a semantic vector, it is possible to unify the processing other than the conversion processing with the input of text data, thereby making it possible to obtain simple output data for a variety of data without complicating the processing.

[0058] The input data may also include image data, which may include at least one of photographic data, pictorial data, graphic data, and text data. That is, the image data may be a photograph of a character string such as a sentence, or a non-textual drawing, or a photograph or painting of a person, object, landscape, etc. Since the image represented by the image data may contain a large amount of information, the input data may not only recognize the name of the main object, but also perform detailed recognition including multiple meanings to convert the information into text from multiple sentences or phrases, such as the facial expression of the object (not limited to animals), the content, size, color, etc. of the expected movement, the positional relationship, size relationship, etc. of multiple objects, and the relationship with the background.

[0059] The input data may also include audio data, which may include at least one of speech and conversation. That is, even non-transcribed audio data may be used as input data. For example, a list of appropriate phrase data can be obtained without the need for the user to organize the main points by generating text in advance from recorded data of a meeting conversation.

[0060] In addition, the information processing method of this embodiment includes an input step of accepting input data having multiple elements, an acquisition step of acquiring a first semantic vector having a direction and indicating the content of the input data based on the direction, a calculation step of converting the acquired first semantic vector into a second semantic vector having a direction different from that of the first semantic vector, and a generation step of extracting multiple words included in the content of the second semantic vector to generate output data. This information processing method eliminates the need to convert output content expressed as semantic vectors into sentences in natural language, and outputs it as a list of phrases, which is time-saving and prevents the user from being presented with unclear content. Therefore, the information processing method can obtain output data that is easy for the user to interpret.

[0061] Furthermore, by having a computer (CPU 11) execute the program 121 of this embodiment, when outputting the content of a semantic vector representing content containing multiple meanings, such as a document, it is possible to easily obtain output that is easy for users to understand through software processing without using special or dedicated hardware, and therefore the program can be easily used by a wide variety of users even on general-purpose PCs, etc.

[0062] The present invention is not limited to the above-described embodiment, and various modifications are possible. For example, in the above embodiment, doc2vec is used for conversion to semantic vectors, but this is not limiting. Any method that can convert vectors into numerical values ​​according to meaning may be used.

[0063] Furthermore, in the above embodiment, the input data and the textualized version of the input data are described as data that can be organized into sentences, but the input data text data, speech content of the audio data, and character strings captured as image data do not need to be organized into sentences. It is also possible to use memos that contain multiple bullet points or in which punctuation marks are not used accurately. Furthermore, data that combines sentences acquired from multiple sources may also be used as input data.

[0064] Furthermore, the order in which each word appears in the output data is not limited to the order in which it appears when the second semantic vector is treated as a series of text. The order may be determined simply in alphabetical order, or if the output includes adjectives and verbs in addition to nouns, the output order may be organized by part of speech.

[0065] In the above embodiment, the output is limited to words and phrases, but this is not necessarily the case. If no processing such as constructing sentences is performed, words that were originally connected and output may be connected and output as they are.

[0066] In the above embodiment, the non-text data is converted to text data before semantic vectors are acquired. However, if semantic vectors can be acquired directly from the non-text data without using text data, conversion to text data is not necessary. Furthermore, conversion to text data is not limited to display content. For example, in the case of image data of a famous painting, the title of the painting and the name of its artist may be used as the text content.

[0067] In the above embodiment, the information processing system 1 is described as being separate from the terminal device 30 that performs input and output and the information processing device 10 that performs processing related to semantic vectors, but this is not limiting. All processing may be performed by a single information processing device (such as a PC). Alternatively, the processing may be further divided and distributed among multiple computers.

[0068] In addition, although the above embodiment has been described assuming processing as an idea generation support operation, the present invention is not limited to this. For example, the present invention may be used to organize key points according to input content, or to search for words that correspond to or are unrelated to an input topic.

[0069] Furthermore, the input and output text does not have to be in Japanese, and the language of the input text and the language of the output text may be different.

[0070] In the above description, the storage unit 12 is described as an example of a computer-readable medium for storing the conversion output control program 121 for controlling the generation of output data for phrases based on semantic vectors of the present invention, which is composed of a hard disk drive (HDD) or a nonvolatile memory such as a flash memory. However, the present invention is not limited to this. Other computer-readable media may include other nonvolatile memories such as MRAM, and portable storage media such as CD-ROMs and DVD discs. Furthermore, a carrier wave may also be used as a medium for providing program data according to the present invention via a communication line. In addition, the specific configurations, contents and procedures of the processing operations, etc. shown in the above embodiments can be modified as appropriate without departing from the spirit of the present invention. The scope of the present invention includes the scope of the invention described in the claims and its equivalents. [Explanation of symbols]

[0071] 1. Information Processing Systems 10. Information processing equipment 11 Control section 12 Storage section 121 Programs 1211 Digitization Processing Unit 1212 Text Conversion Unit 13 Communications Department 20 Database Device 21 Memory section 211 Terminology Data 30 Terminal Equipment 31 Control Unit 32 Storage section 33 Display section 34 Operation reception section 35 Communications Department

Claims

1. an input means for accepting input data having a plurality of first elements; an acquisition means for acquiring a first semantic vector having a direction and indicating the content of the input data by the direction; a computing means for converting the acquired first semantic vector into a single second semantic vector having a different direction from the first semantic vector; a generating means for generating output data by converting the single second semantic vector back into a string of text having a plurality of second elements, dividing the text into a plurality of words, and extracting a plurality of meaningful phrases obtained based on the words as the plurality of second elements; An information processing system comprising:

2. The series of texts is an incomplete sentence.

2. The information processing system according to claim 1.

3. 3. The information processing system according to claim 1, wherein the plurality of phrases are words or phrases.

4. 4. The information processing system according to claim 1, wherein the input data includes at least one of image data, text data, and voice data.

5. the input data includes text data; The text data is sentence data.

5. The information processing system according to claim 4.

6. the input data includes non-text data; a conversion means for converting the non-text data into text data including the plurality of first elements; 6. The information processing system according to claim 1, wherein:

7. the input data includes image data; The image data includes at least one of photograph data, picture data, graphic data, and character data.

7. The information processing system according to claim 4, wherein:

8. the input data includes audio data; The audio data includes at least one of speech and conversation data.

7. The information processing system according to claim 4, wherein:

9. An information processing method by a control unit of a computer, comprising: an input step of accepting input data having a plurality of first elements; an acquisition step of acquiring one first semantic vector having a direction and indicating the content of the input data by the direction; a calculation step of converting the acquired first semantic vector into a single second semantic vector having a different direction from the first semantic vector; a generation step of inversely converting the single second semantic vector into a string of text having a plurality of second elements, dividing the text into a plurality of words, and extracting a plurality of meaningful phrases obtained based on the words as the plurality of second elements to generate output data; An information processing method comprising:

10. Computer an input means for accepting input data having a plurality of first elements; an acquisition means for acquiring a first semantic vector having a direction and indicating the content of the input data by the direction; a computing means for converting the acquired first semantic vector into a single second semantic vector having a different direction from the first semantic vector; a generating means for converting the single second semantic vector back into a string of text having a plurality of second elements, dividing the text into a plurality of words, and extracting a plurality of meaningful phrases obtained based on the words as the plurality of second elements to generate output data; A program characterized by functioning as

Citation Information

Patent Citations

  • Learning apparatus, generation device, learning method, generation method, learning program, generation program, and model

    JP2019057034A

  • Idea support apparatus and program

    JP2019086815A

  • Idea support device, idea support system and program

    JP2020197957A