Information processing device, information processing method, and information processing program

The information processing apparatus and method improve digest video generation by accurately selecting video sections based on similarity with inclusion and exclusion prompts, ensuring relevant content is included and irrelevant content is excluded.

WO2025158531A1PCT designated stage expired Publication Date: 2025-07-31NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/001843
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing techniques for generating digest videos inaccurately include video sections that should be excluded, leading to unsuitable content extraction.

Method used

An information processing apparatus and method that utilizes image acquisition, prompt acquisition, and similarity calculation units to determine the degree of similarity between image data and prompts representing events to be included or excluded, calculating a score for inclusion in the digest based on these similarities.

Benefits of technology

Generates a digest video more suitable for the required extraction content by accurately including or excluding relevant sections, enhancing decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024001843_31072025_PF_FP_ABST
    Figure JP2024001843_31072025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device is provided with: a first similarity degree acquisition unit that acquires a first degree of similarity indicating the degree of similarity between a feature value of image data and a feature value of a first prompt representing an event that is required to be included in a digest, the first degree of similarity being obtained by inputting the image data and the first prompt to a calculator; a second similarity degree acquisition unit that acquires a second degree of similarity indicating the degree of similarity between the feature value of the image data and a feature value of a second prompt representing an event that is required not to be included in the digest, the second degree of similarity being obtained by inputting the image data and the second prompt to the calculator; and a score calculation unit that uses the first degree of similarity and the second degree of similarity to calculate a score for whether to include the image data in the digest.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and information processing program

[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing program.

[0002] Techniques for generating digest videos from videos are known. One example of a technique for generating digest videos is described in Patent Document 1. The summary video generation device described in Patent Document 1 includes a feature vector calculation unit that calculates, for each video segment, feature vectors of multiple modals for multiple time scales including the video segment; a video segment importance calculation unit that calculates the importance of the video segment from the multiple feature vectors for each video segment using a pre-trained neural network; and a video summarization unit that extracts video segments from the video to be summarized in descending order of importance up to a predetermined total length, connects the extracted video segments in chronological order, and generates a summary video.

[0003] Japanese Patent Application Publication No. 2023-122672

[0004] In the technique described in Patent Document 1, there are cases where the importance of a video section that is not desired to be included in the digest video is calculated to be high.

[0005] The present disclosure has been made in view of the above-mentioned problems, and an exemplary purpose thereof is to provide a technique for generating a digest that is more suitable for the requested extracted content.

[0006] An information processing device according to an exemplary aspect of the present disclosure includes an image acquisition means for acquiring image data, a prompt acquisition means for acquiring a first prompt representing an event that is requested to be included in a digest and a second prompt representing an event that is requested not to be included in the digest, a first similarity acquisition means for acquiring a first similarity indicating a degree of similarity between a feature of the image data and a feature of the first prompt by inputting the image data and the first prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt, a second similarity acquisition means for acquiring a second similarity indicating a degree of similarity between a feature of the image data and a feature of the second prompt by inputting the image data and the second prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt, and a score calculation means for calculating a score regarding whether to include the image data in the digest using the first similarity and the second similarity.

[0007] An information processing method according to an exemplary aspect of the present disclosure includes: an image acquisition process in which at least one processor acquires image data; a prompt acquisition process in which the at least one processor acquires a first prompt representing an event to be included in a digest and a second prompt representing an event to be excluded from the digest; a first similarity acquisition process in which the at least one processor inputs the image data and the first prompt into a calculator that calculates the similarity between the feature amounts of the image data and the feature amounts of the prompt, and acquires a first similarity indicating a degree of similarity between a feature amount of the image data and a feature amount of the first prompt; a second similarity acquisition process in which the at least one processor inputs the image data and the second prompt into a calculator that calculates the similarity between the feature amounts of the image data and the feature amounts of the prompt, and acquires a second similarity indicating a degree of similarity between a feature amount of the image data and a feature amount of the second prompt; and a score calculation process in which the at least one processor uses the first similarity and the second similarity to calculate a score for whether to include the image data in the digest.

[0008] An information processing program according to an exemplary aspect of the present disclosure causes a computer to function as an image acquisition means for acquiring image data, a prompt acquisition means for acquiring a first prompt representing an event that is requested to be included in a digest and a second prompt representing an event that is requested not to be included in the digest, a first similarity acquisition means for acquiring a first similarity indicating a degree of similarity between a feature of the image data and a feature of the first prompt by inputting the image data and the first prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt, a second similarity acquisition means for acquiring a second similarity indicating a degree of similarity between a feature of the image data and a feature of the second prompt by inputting the image data and the second prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt, and a score calculation means for calculating a score regarding whether to include the image data in the digest using the first similarity and the second similarity.

[0009] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that a technique for generating a digest that is more suitable for the requested extracted content can be provided.

[0010] FIG. 1 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 2 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 3 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 4 is a block diagram showing an example of the functional configuration of a control unit of an information processing device according to the present disclosure. FIG. 5 is a diagram showing a specific example of a summary score calculation process. FIG. 6 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 7 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 8 is a block diagram showing an example of the functional configuration of an information processing device according to the present disclosure. FIG. 9 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 10 is a block diagram showing the configuration of a computer functioning as an information processing device according to the present disclosure.

[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0012] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0013] (Configuration of information processing device) The configuration of the information processing device 1 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1. As shown in Fig. 1, the information processing device 1 includes an image acquisition unit 11, a prompt acquisition unit 12, a first similarity acquisition unit 13, a second similarity acquisition unit 14, and a score calculation unit 15.

[0014] The image acquisition unit 11 acquires image data. The prompt acquisition unit 12 acquires a first prompt representing an event to be included in the digest and a second prompt representing an event not to be included in the digest. The first similarity acquisition unit 13 acquires a first similarity indicating the degree of similarity between a feature of the image data and a feature of the first prompt, obtained by inputting the image data and the first prompt into a calculator that calculates the similarity between a feature of the image data and a feature of the prompt. The second similarity acquisition unit 14 acquires a second similarity indicating the degree of similarity between a feature of the image data and a feature of the second prompt, obtained by inputting the image data and the second prompt into a calculator that calculates the similarity between a feature of the image data and a feature of the prompt. The score calculation unit 15 calculates a score for whether to include the image data in the digest using the first similarity and the second similarity.

[0015] (Effects of Information Processing Device) As described above, the information processing device 1 employs a configuration including an image acquisition unit 11 that acquires image data, a prompt acquisition unit 12 that acquires a first prompt representing an event that is requested to be included in the digest and a second prompt representing an event that is requested not to be included in the digest, a first similarity acquisition unit 13 that acquires a first similarity indicating a degree of similarity between a feature amount of the image data and a feature amount of the first prompt by inputting the image data and the first prompt into a calculator that calculates the similarity between a feature amount of the image data and a feature amount of the prompt, a second similarity acquisition unit 14 that acquires a second similarity indicating a degree of similarity between a feature amount of the image data and a feature amount of the second prompt by inputting the image data and the second prompt into a calculator that calculates the similarity between the feature amount of the image data and a feature amount of the prompt, and a score calculation unit 15 that calculates a score for whether to include the image data in the digest using the first similarity and the second similarity. Therefore, the information processing device 1 has the effect of being able to calculate a score for generating a digest that is more suitable for the requested extracted content.

[0016] (Flow of Information Processing Method) The flow of information processing method S1 will be described with reference to Fig. 2. Fig. 2 is a flowchart showing the flow of information processing method S1. As shown in Fig. 2, information processing method S1 includes an image acquisition process S11, a prompt acquisition process S12, a first similarity acquisition process S13, a second similarity acquisition process S14, and a score calculation process S15.

[0017] In an image acquisition process S11, at least one processor acquires image data. In a prompt acquisition process S12, the at least one processor acquires a first prompt representing an event desired to be included in the digest and a second prompt representing an event desired not to be included in the digest.

[0018] In the first similarity acquisition process S13, the at least one processor acquires a first similarity indicating the degree of similarity between the features of the image data and the features of the first prompt, which is obtained by inputting the image data and the first prompt into a calculator that calculates the similarity between the features of the image data and the features of the prompt.

[0019] In the second similarity acquisition process S14, the at least one processor acquires a second similarity indicating the degree of similarity between the features of the image data and the features of the second prompt, which is obtained by inputting the image data and the second prompt into a calculator that calculates the similarity between the features of the image data and the features of the prompt.

[0020] In the score calculation process S15, the at least one processor uses the first similarity and the second similarity to calculate a score as to whether the image data should be included in the digest.

[0021] (Effects of Information Processing Method) As described above, the information processing method S1 includes an image acquisition process S11 in which at least one processor acquires image data, a prompt acquisition process S12 in which the at least one processor acquires a first prompt representing an event that is requested to be included in a digest and a second prompt representing an event that is requested not to be included in the digest, and a calculator for calculating a degree of similarity between a feature amount of the image data and a feature amount of the prompt, the degree of similarity being obtained by inputting the image data and the first prompt into a calculator for calculating a degree of similarity between a feature amount of the image data and a feature amount of the prompt. a first similarity acquisition process S13 for acquiring a first similarity indicating the degree of similarity between the features of the image data and the features of the second prompt, a second similarity acquisition process S14 for acquiring a second similarity indicating the degree of similarity between the features of the image data and the features of the second prompt by inputting the image data and the second prompt into a calculator that calculates the similarity between the features of the image data and the features of the second prompt, and a score calculation process S15 for calculating a score for whether to include the image data in the digest using the first similarity and the second similarity. Thus, the information processing method S1 has the effect of being able to calculate a score for generating a digest that is more suitable for the requested extraction content.

[0022] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0023] (Configuration of Information Processing Device) The information processing device 1A according to the present disclosure is a device that calculates a score for determining whether image data should be included in a digest. The image data is data representing a still image or a moving image. The image data is, for example, data representing an image captured by a camera, but is not limited to this. The digest is image data representing a summary of a video. For example, the digest is video data including video data of multiple sections extracted from video data of a baseball game.

[0024] The configuration of the information processing device 1A will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of the information processing device 1A. The information processing device 1A includes a control unit 10A, a storage unit 20A, a communication unit 30A, an input unit 40A, and an output unit 50A.

[0025] (Communication Unit) The communication unit 30A communicates with devices external to the information processing device 1A via a communication line. While the specific configuration of the communication line does not limit the present exemplary embodiment, examples of the communication line include a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination thereof. The communication unit 30A transmits data supplied from the control unit 10A to other devices, and supplies data received from other devices to the control unit 10A.

[0026] (Input Unit) The input unit 40A is configured to receive input to the information processing device 1A, and includes, for example, input devices such as a keyboard, a mouse, a touch panel, a camera, a microphone, etc. The input unit 40A may also be configured to receive data from the input devices via an interface such as a USB (Universal Serial Bus).

[0027] (Output Unit) The output unit 50A is a component for performing output from the information processing device 1A, and includes, for example, output devices such as a display, a printer, a touch panel, a speaker, etc. The output unit 50A may be configured to include, for example, an interface such as a USB, and to output data to the output device via the interface.

[0028] (Storage Unit) The storage unit 20A stores various types of information referenced by the control unit 10A. Examples of such information include image data 201, positive prompt 202, negative prompt 203, first encoder 211, and second encoder 212. Note that the first encoder 211 and the second encoder 212 being stored in the storage unit 20A means that parameters defining the first encoder 211 and parameters defining the second encoder 212 are stored in the storage unit 20A. The first encoder 211 and the second encoder 212 are examples of a calculator according to the present disclosure. The image data 201 is image data to be determined as to whether to include it in a digest.

[0029] (Positive Prompt) The positive prompt 202 is a prompt that represents an event that is requested to be included in the digest. The positive prompt 202 is an example of a first prompt according to the present disclosure. The positive prompt 202 includes, for example, at least one of text and image data. If the image data 201 is a photographed image of a baseball game, the positive prompt 202 may include, for example, the text "Hitters who hit home runs, pitcher throws the ball."

[0030] (Negative Prompt) The negative prompt 203 is a prompt that represents an event that is not to be included in the digest. The negative prompt is an example of a second prompt according to the present disclosure. For example, the negative prompt 203 includes at least one of text and image data. If the image data 201 is a photographed image of a baseball game, the negative prompt 203 may include, for example, the text "players sitting on the bench, sky."

[0031] (First Encoder and Second Encoder) The first encoder 211 and the second encoder 212 are calculators for calculating the similarity between the feature amounts of image data and the feature amounts of prompts. The first encoder 211 is used by the positive score calculation unit 14A (described later) to calculate positive scores. The second encoder 212 is used by the negative score calculation unit 15A (described later) to calculate negative scores.

[0032] The first encoder 211 and the second encoder 212 each receive image data and a prompt, and output a similarity between the image data and the prompt. Examples of the first encoder 211 and the second encoder 212 include a visual language model (VLM) generated by machine learning.

[0033] (Control Unit) FIG. 4 is a diagram illustrating an example of the functional configuration of the control unit 10A. The control unit 10A includes an image input unit 11A, a positive prompt input unit 12A, a negative prompt input unit 13A, a positive score calculation unit 14A, a negative score calculation unit 15A, a summary score calculation unit 16A, a score output unit 17A, and a digest generation unit 18A. The image input unit 11A is an example of an image acquisition means according to the present disclosure. The positive prompt input unit 12A and the negative prompt input unit 13A are examples of a prompt acquisition means according to the present disclosure. The positive score calculation unit 14A is an example of a first similarity acquisition means according to the present disclosure. The negative score calculation unit 15A is an example of a second similarity acquisition means according to the present disclosure. The summary score calculation unit 16A is an example of a score calculation means according to the present disclosure. The score output unit 17A is an example of a score output means according to the present disclosure. The digest generation unit 18A is an example of a digest generation means according to the present disclosure.

[0034] (Image Input Unit) The image input unit 11A acquires image data 201 and supplies the acquired image data to the positive score calculation unit 14A and the negative score calculation unit 15A. As an example, the image input unit 11A may acquire the image data 201 by reading the image data 201 from a storage destination (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A) specified by the user of the information processing device 1A. The image input unit 11A may also acquire the image data 201 by receiving the image data 201 from another device via the communication unit 30A. The image input unit 11A may also acquire the image data 201 input to the input unit 40A.

[0035] (Positive prompt input unit) The positive prompt input unit 12A acquires a positive prompt 202 and supplies the acquired positive prompt 202 to the positive score calculation unit 14A. As an example, the positive prompt input unit 12A may acquire the positive prompt 202 by reading the positive prompt 202 from a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A). Alternatively, the positive prompt input unit 12A may acquire the positive prompt 202 by receiving the positive prompt 202 from another device via the communication unit 30A. Alternatively, the positive prompt input unit 12A may acquire the positive prompt 202 input to the input unit 40A.

[0036] (Negative Prompt Input Unit) The negative prompt input unit 13A acquires the negative prompt 203 and supplies the acquired negative prompt 203 to the negative score calculation unit 15A. As an example, the negative prompt input unit 13A may acquire the negative prompt 203 by reading the negative prompt 203 from a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A). The negative prompt input unit 13A may also acquire the negative prompt 203 by receiving the negative prompt 203 from another device via the communication unit 30A. The negative prompt input unit 13A may also acquire the negative prompt 203 input to the input unit 40A.

[0037] (Positive score calculation unit) The positive score calculation unit 14A obtains a positive score indicating the degree of similarity between the image data 201 and the positive prompt 202. The positive score is an example of a first similarity according to the present disclosure. As an example, a larger positive score indicates a higher degree of recommendation for inclusion in the digest. The positive score calculation unit 14A calculates the positive score using the first encoder 211. In other words, the positive score calculation unit 14A inputs the image data 201 and the positive prompt 202 into the first encoder 211 stored in the memory unit 20A, and obtains the positive score that is its output.

[0038] As shown in FIG. 4, the positive score calculation unit 14A includes a first feature extraction unit 141A and a first distance calculation unit 142A. The first feature extraction unit 141A and the first distance calculation unit 142A are realized by the positive score calculation unit 14A and the first encoder 211. The first feature extraction unit 141A extracts features from the image data 201 and the positive prompt 202. The first distance calculation unit 142A calculates a positive score indicating the distance between the feature of the image data 201 and the feature of the positive prompt 202. As an example, the larger the positive score value, the higher the similarity between the image data 201 and the positive prompt. As an example, the positive score may be a real number between 0 and 1.

[0039] (Negative Score Calculation Unit) The negative score calculation unit 15A acquires a negative score indicating the degree of similarity between the image data 201 and the negative prompt 203. As an example, a larger negative score indicates a higher degree of recommendation not to include the image data in the digest. The negative score is an example of a first similarity according to the present disclosure. The negative score calculation unit 15A calculates the negative score using the second encoder 212. In other words, the negative score calculation unit 15A inputs the image data 201 and the negative prompt 203 into the second encoder 212 stored in the memory unit 20A, and acquires the negative score that is the output.

[0040] The negative score calculation unit 15A includes a second feature extraction unit 151A and a second distance calculation unit 152A. The second feature extraction unit 151A and the second distance calculation unit 152A are realized by the negative score calculation unit 15A and the second encoder 212. The second feature extraction unit 151A extracts features from the image data 201 and the negative prompt 203, respectively. The second distance calculation unit 152A calculates a negative score indicating the distance between the feature of the image data 201 and the feature of the negative prompt 203. As an example, the larger the negative score value, the higher the similarity between the image data 201 and the negative prompt. As an example, the negative score may be a real number between 0 and 1.

[0041] The positive score calculation unit 14A and the negative score calculation unit 15A may calculate the scores using a common encoder. In other words, the encoder used by the positive score calculation unit 14A may be the same as or different from the encoder used by the negative score calculation unit 15A.

[0042] (Summary Score Calculation Unit) The summary score calculation unit 16A uses the positive score and the negative score to calculate a summary score for whether to include the image data 201 in the digest. As an example, the higher the summary score, the higher the degree to which it is recommended to include the image data in the digest. As an example, the summary score calculation unit 16A calculates the summary score by subtracting the negative score from the positive score.

[0043] (Score Output Unit) The score output unit 17A outputs the summary score calculated by the summary score calculation unit 16A. For example, the score output unit 17A may output the summary score by writing it to a storage destination (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A) designated by the user of the information processing device 1A. Furthermore, the score output unit 17A may output the summary score by transmitting the summary score via the communication unit 30A, or may output the summary score to an output device such as a display.

[0044] (Digest Generation Unit) The digest generation unit 18A generates a digest including image data whose summary scores calculated by the summary score calculation unit 16A satisfy a predetermined condition. As an example, the digest generation unit 18A may extract image data whose summary scores are greater than a threshold from the image data and generate a digest including the extracted image data. Alternatively, the digest generation unit 18A may sort the image data in ascending order of summary score, count from the top, extract image data that are included in a predetermined percentage, and generate a digest including the extracted image data.

[0045] The digest generation unit 18A outputs the generated digest. As an example, the digest generation unit 18A may output the digest by writing the digest to a storage destination (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A) designated by the user of the information processing device 1A. Furthermore, the score output unit 17A may output the digest by transmitting the digest via the communication unit 30A, or may output the digest to an output device such as a display.

[0046] 5 is a diagram showing a specific example of the process of calculating a summary score by the summary score calculation unit 16A. In the example of FIG. 5, image data 201 and a positive prompt 202 are input to a first encoder 211, and image data 201 and a negative prompt 203 are input to a second encoder 212. In addition, a positive score of "0.8" output by the first encoder 211 and a negative score of "0.2" output by the second encoder 212 are input to the summary score calculation unit 16A. The summary score calculation unit 16A outputs a value of "0.6" obtained by subtracting the negative score of "0.2" from the positive score of "0.6" as the summary score.

[0047] (Flow of Information Processing Method) Fig. 6 is a flow diagram showing an example of the flow of an information processing method S1A executed by the information processing device 1A. Some of the steps included in the flow diagram of Fig. 6 may be executed in parallel or in a different order.

[0048] In step S101, the image input unit 11A acquires image data 201, the positive prompt input unit 12A acquires a positive prompt, and the negative prompt input unit 13A acquires a negative prompt. In step S101, the image input unit 11A may acquire image data for each frame of a plurality of frames constituting a moving image, or may acquire image data for a plurality of consecutive frames in a single step.

[0049] In step S102, the first feature extraction unit 141A extracts the feature of the image data 201 and the feature of the positive prompt 202, and the second feature extraction unit 151A extracts the feature of the image data 201 and the feature of the negative prompt 203.

[0050] In step S103, the first distance calculation unit 142A calculates the positive scores, and the second distance calculation unit 152A calculates the negative scores. In step S104, the summary score calculation unit 16A calculates the summary scores using the negative scores and the positive scores. In step S105, the score output unit 17A outputs the summary scores.

[0051] In step S106, the score output unit 17A determines whether to end the score calculation process. As an example, if image data for which the score is to be calculated remains, the score output unit 17A determines not to end the calculation process. On the other hand, if image data for which the score is to be calculated does not remain, the score output unit 17A determines to end the calculation process. If the score output unit 17A determines to end the calculation process (YES in step S106), the score output unit 17A proceeds to the processing of step S107. On the other hand, if the score output unit 17A determines not to end the calculation process (NO in step S106), the score output unit 17A returns to the processing of step S101. In step S107, the digest generation unit 18A generates a digest based on the summary score.

[0052] (Use Case) The information processing device 1A according to the present disclosure can be utilized in various fields. For example, the information processing device 1A according to the present disclosure can be utilized in the medical / healthcare field. For example, it is conceivable to generate a digest by extracting a video of a person performing a specific action from a video taken in a medical facility such as a hospital. For example, when generating a digest from a video of a person undergoing rehabilitation, the information processing device 1A calculates a summary score using a positive prompt such as "a person walking while holding onto a handrail" and a negative prompt such as "a person sitting." This prevents the summary score of a video of a person sitting from being too high, and allows for more appropriate calculation of the summary score.

[0053] In this example, the image data is data obtained by photographing a subject. The digest may also be used to support decision-making regarding the subject. For example, the subject may be, for example, a patient, medical staff (including medical professionals), or visitors. For example, a medical professional may review the generated digest to determine whether a patient is performing rehabilitation appropriately and encourage rehabilitation for patients who are not performing rehabilitation appropriately.

[0054] Furthermore, the information processing device 1A may calculate the summary score using, for example, a positive prompt such as "person who fell" and a negative prompt such as "child." In this case, it is possible to prevent a video of a child lying down playfully from having a high summary score, thereby allowing for more appropriate calculation of the summary score. Furthermore, for example, by checking the generated digest, medical personnel can identify patients who have fallen within a medical facility, thereby enabling them to receive appropriate treatment.

[0055] As described above, the information processing device 1A is configured to generate a digest that includes image data for which the score calculated by the summary score calculation unit 16A satisfies a predetermined condition. Therefore, the information processing device 1A has the effect of being able to generate a digest that is more suitable for the required extraction condition.

[0056] Furthermore, the information processing device 1A employs a configuration in which the image data 201 is data obtained by photographing the subject, and the digest is data representing an image used to support decision-making regarding the subject. Therefore, the information processing device 1A can achieve the effect of more effectively supporting decision-making regarding the subject by using a digest that is more suited to the required extraction conditions.

[0057] Furthermore, the information processing device 1A employs a configuration in which the calculator that calculates the similarity between the feature amounts of image data and the feature amounts of the prompt is a visual language model generated by machine learning. Therefore, the information processing device 1A can calculate a more appropriate summary score by using the positive score and negative score calculated using the visual language model generated by machine learning.

[0058] Furthermore, the information processing device 1A employs a configuration in which the positive prompts 202 and the negative prompts 203 each include at least one of text and image data. Therefore, the information processing device 1A can more appropriately calculate the summary score by using positive prompts and negative prompts that include text and / or image data.

[0059] The information processing device 1A also includes a score output unit 17A that outputs the summary score calculated by the summary score calculation unit 16A. Therefore, the information processing device 1A allows a user of the information processing device 1A to understand the calculated summary score.

[0060] [Third Exemplary Embodiment] A third exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0061] (Configuration of Information Processing Apparatus) Fig. 7 is a block diagram showing the configuration of an information processing apparatus 1B. The information processing apparatus 1A includes a control unit 10B, a storage unit 20B, a communication unit 30A, an input unit 40A, and an output unit 50A.

[0062] In addition to image data 201, positive prompt 202, negative prompt 203, first encoder 211, and second encoder 212, storage unit 20B also stores must prompt 204, never prompt 205, third encoder 213, and fourth encoder 214. Note that storing third encoder 213 and fourth encoder 214 in storage unit 20B means that parameters defining third encoder 213 and parameters defining fourth encoder 214 are stored in storage unit 20B. Third encoder 213 and fourth encoder 214 are examples of calculators according to the present disclosure.

[0063] Must prompts 204 are prompts that represent events that must be included in the digest, while never prompts 205 are prompts that represent events that must not be included in the digest. Each of must prompts 204 and never prompts 205 includes, for example, text and / or image data.

[0064] The third encoder 213 and the fourth encoder 214 are calculators for calculating the similarity between the feature amounts of the image data and the feature amounts of the prompt. The third encoder 213 is used by a must-score calculation unit 24B (described later) to calculate a must score. The fourth encoder 214 is used by a never-score calculation unit 25B (described later) to calculate a never score.

[0065] The third encoder 213 and the fourth encoder 214 each receive image data and a prompt, and output a similarity between the image data and the prompt. Examples of the third encoder 213 and the fourth encoder 214 include a visual language model (VLM) generated by machine learning.

[0066] 8 is a diagram schematically illustrating an example of the functional configuration of the control unit 10B. The control unit 10B includes an image input unit 11A, a positive prompt input unit 12A, a negative prompt input unit 13A, a positive score calculation unit 14A, a negative score calculation unit 15A, a score output unit 17A, and a digest generation unit 18A, as well as a must prompt input unit 22B, a never prompt input unit 23B, a must score calculation unit 24B, a never score calculation unit 25B, and a summary score calculation unit 16B. The must prompt input unit 22B and the never prompt input unit 23B are examples of prompt acquisition means according to the present disclosure. The must score calculation unit 24B is an example of third similarity acquisition means according to the present disclosure. The never score calculation unit 25B is an example of fourth similarity acquisition means according to the present disclosure.

[0067] The must prompt input unit 22B acquires the must prompt 204 and supplies the acquired must prompt 204 to the must score calculation unit 24B. As an example, the must prompt input unit 22B may acquire the must prompt 204 by reading it from a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A). Alternatively, the must prompt input unit 22B may acquire the must prompt 204 by receiving it from another device via the communication unit 30A. Alternatively, the must prompt input unit 22B may acquire the must prompt 204 input to the input unit 40A.

[0068] The never prompt input unit 23B acquires the never prompt 205 and supplies the acquired never prompt 205 to the never score calculation unit 25B. As an example, the never prompt input unit 23B may acquire the never prompt 205 by reading it from a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A). The never prompt input unit 23B may also acquire the never prompt 205 by receiving it from another device via the communication unit 30A. The never prompt input unit 23B may also acquire the never prompt 205 input to the input unit 40A.

[0069] The must score calculation unit 24B acquires a must score indicating the degree of similarity between the image data 201 and the must prompt 204. The must score is an example of a third similarity according to the present disclosure. As an example, a larger must score indicates a higher likelihood of inclusion in the digest. The must score calculation unit 24B calculates the must score using the third encoder 213. In other words, the must score calculation unit 24B acquires a must score indicating the degree of similarity between the feature amounts of the image data 201 and the feature amounts of the must prompt 204, which is obtained by inputting the image data 201 and the must prompt 204 into the third encoder 213, which calculates the similarity between the feature amounts of the image data and the feature amounts of the prompt.

[0070] As shown in FIG. 8 , the must score calculation unit 24B includes a third feature extraction unit 241B and a third distance calculation unit 242B. The third feature extraction unit 241B and the third distance calculation unit 242B are realized by the must score calculation unit 24B and the third encoder 213. The third feature extraction unit 241B extracts features from the image data 201 and the must prompt 204, respectively. The third distance calculation unit 242B calculates a must score that indicates the distance between the feature of the image data 201 and the feature of the must prompt 204. As an example, the larger the must score value, the higher the similarity between the image data 201 and the must prompt. As an example, the must score may be a real number between 0 and 1.

[0071] (Never score calculation unit) The never score calculation unit 25B acquires a never score indicating the degree of similarity between the image data 201 and the never prompt 205. The never score is an example of a fourth similarity according to the present disclosure. As an example, a larger never score indicates a higher degree of not to be included in the digest. The never score calculation unit 25B calculates the never score using the fourth encoder 214. In other words, the never score calculation unit 25B acquires a never score indicating the degree of similarity between the feature amounts of the image data 201 and the feature amounts of the never prompt 205, which is obtained by inputting the image data 201 and the never prompt 205 into the fourth encoder 214, which calculates the similarity between the feature amounts of the image data and the feature amounts of the prompt.

[0072] As shown in FIG. 8 , the never score calculation unit 25B includes a fourth feature extraction unit 251B and a fourth distance calculation unit 252B. The fourth feature extraction unit 251B and the fourth distance calculation unit 252B are realized by the never score calculation unit 25B and the fourth encoder 214. The fourth feature extraction unit 251B extracts features from the image data 201 and the never prompt 205. The fourth distance calculation unit 252B calculates a never score indicating the distance between the feature of the image data 201 and the feature of the never prompt 205. As an example, the never score indicates that the larger the value, the higher the similarity between the image data 201 and the never prompt. As an example, the never score may be a real number between 0 and 1.

[0073] The summary score calculation unit 16B calculates the summary score using the must score and the never score in addition to the positive score and the negative score. As an example, if the must score is equal to or greater than a predetermined threshold, the summary score calculation unit 16B may calculate a summary score indicating that the image data 201 should be included in the digest. Also, as an example, if the never score is equal to or greater than a predetermined threshold, the summary score calculation unit 16B may calculate a summary score indicating that the image data 201 should not be included in the digest.

[0074] (Flow of Information Processing Method) Fig. 9 is a flow diagram showing an example of the flow of an information processing method S1B executed by information processing device 1B. Some of the steps included in the flow diagram of Fig. 9 may be executed in parallel or in a different order.

[0075] In step S201, the image input unit 11A acquires image data 201, the positive prompt input unit 12A acquires a positive prompt, the negative prompt input unit 13A acquires a negative prompt, the must prompt input unit 22B acquires a must prompt, and the never prompt input unit 23B acquires a never prompt.

[0076] In step S202, the first feature extraction unit 141A extracts the feature of the image data 201 and the feature of the positive prompt 202, and the second feature extraction unit 151A extracts the feature of the image data 201 and the feature of the negative prompt 203. Furthermore, the third feature extraction unit 241B extracts the feature of the image data 201 and the feature of the must prompt 204, and the fourth feature extraction unit 251B extracts the feature of the image data 201 and the feature of the never prompt 205.

[0077] In step S203, the first distance calculation unit 142A calculates the positive score, and the second distance calculation unit 152A calculates the negative score. Furthermore, the third distance calculation unit 242B calculates the must score, and the fourth distance calculation unit 252B calculates the never score. In step S204, the summary score calculation unit 16B calculates the summary score using the negative score, positive score, must score, and never score. In step S205, the score output unit 17A outputs the summary score.

[0078] In step S206, the score output unit 17A determines whether to end the score calculation process. As an example, if image data for which the score is to be calculated remains, the score output unit 17A determines not to end the calculation process. On the other hand, if image data for which the score is to be calculated no longer remains, the score output unit 17A determines to end the calculation process. If the score output unit 17A determines to end the calculation process (YES in step S206), the score output unit 17A proceeds to the processing of step S207. On the other hand, if the score output unit 17A determines not to end the calculation process (NO in step S206), the score output unit 17A returns to the processing of step S201. In step S207, the digest generation unit 18A generates a digest based on the summary score.

[0079] (Effects of the Information Processing Device) As described above, in the information processing device 1B, the must prompt input unit 22B acquires a must prompt representing an event that must be included in a digest, and further includes a must score calculation unit 24B that acquires a must score indicating the degree of similarity between the feature amounts of the image data 201 and the feature amounts of the must prompt, which is obtained by inputting the image data 201 and the must prompt to the third encoder 213. The summary score calculation unit 16B is configured to calculate a summary score indicating that the image data 201 should be included in the digest if the must score is equal to or greater than a predetermined threshold. Therefore, the information processing device 1B can include image data similar to the must prompt in a digest, regardless of whether the negative score is high. This has the effect of enabling the generation of a digest that is more suitable for the required extraction conditions.

[0080] Furthermore, in the information processing device 1B, the never prompt input unit 23B acquires a never prompt representing an event that must not be included in the digest, and further includes a never score calculation unit 25B that acquires a never score indicating the degree of similarity between the feature amounts of the image data 201 and the feature amounts of the never prompt 205, obtained by inputting the image data 201 and the never prompt to the fourth encoder 214. The summary score calculation unit 16B is configured to calculate a summary score indicating that the image data 201 should not be included in the digest if the never score is below a predetermined threshold. Therefore, with the information processing device 1B, image data similar to the never prompt can be excluded from the digest regardless of whether the positive score is high. This has the effect of enabling the generation of a digest that is more suitable for the required extraction conditions.

[0081] [Variation] In the exemplary embodiment described above, the positive score calculation unit 14A calculated the positive score using the first encoder 211 stored in the storage unit 20A. The method by which the positive score calculation unit 14A calculates the positive score is not limited to the above example. The positive score calculation unit 14A may calculate the positive score using, for example, a calculator stored in a device other than the information processing device 1A. When the calculator is stored in a device other than the information processing device 1A, the positive score calculation unit 14A, for example, inputs the image data 201 and the positive prompt 202 to the calculator by transmitting the image data 201 and the positive prompt 202 to the device storing the calculator via the communication unit 30A. In this case, the positive score calculation unit 14A obtains the positive score calculated by the calculator by receiving it from the device via the communication unit 30A.

[0082] Furthermore, the positive score calculation unit 14A may input the image data 201 and the positive prompt 202 to the calculator by outputting the image data 201 and the positive prompt 202 to a device that stores the calculator via the output unit 50A. In this case, the positive score calculation unit 14A acquires the positive score calculated by the calculator from the device via the input unit 40A.

[0083] Furthermore, the second encoder 212 may be stored in a device other than the information processing device 1A. In this case, as an example, the negative score calculation unit 15A inputs the image data 201 and the negative prompt 203 to the calculator by transmitting the image data 201 and the negative prompt 203 to a device that stores the calculator via the communication unit 30A. In this case, the negative score calculation unit 15A obtains the negative score calculated by the calculator by receiving it from the above device.

[0084] Furthermore, the negative score calculation unit 15A may input the image data 201 and the negative prompt 203 to the calculator by outputting the image data 201 and the negative prompt 203 to a device that stores the calculator via the output unit 50A. In this case, the negative score calculation unit 15A acquires the negative score calculated by the calculator from the device via the input unit 40A.

[0085] [Example of implementation by software] Some or all of the functions of the information processing devices 1, 1A, 1B (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.

[0086] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 10. Figure 10 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.

[0087] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.

[0088] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0089] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0090] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0091] [Appendix A] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0092] (Appendix A1) An information processing device comprising: an image acquisition means for acquiring image data; a prompt acquisition means for acquiring a first prompt representing an event that is requested to be included in a digest and a second prompt representing an event that is requested not to be included in the digest; a first similarity acquisition means for acquiring a first similarity indicating a degree of similarity between a feature of the image data and a feature of the first prompt, the first similarity being obtained by inputting the image data and the first prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt; a second similarity acquisition means for acquiring a second similarity indicating a degree of similarity between a feature of the image data and a feature of the second prompt, the second similarity being obtained by inputting the image data and the second prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt; and a score calculation means for calculating a score for whether to include the image data in the digest using the first similarity and the second similarity.

[0093] (Appendix A2) The information processing device described in Appendix A1, wherein the prompt acquisition means acquires, in addition to the first prompt and the second prompt, a third prompt representing an event that must be included in the digest, and further comprises third similarity acquisition means for acquiring a third similarity indicating a degree of similarity between features of the image data and features of the third prompt by inputting the image data and the third prompt into a calculator that calculates the similarity between features of the image data and features of the prompt, and the score calculation means calculates a score indicating that the image data should be included in the digest if the third similarity is equal to or greater than a predetermined threshold.

[0094] (Appendix A3) The information processing device according to Appendix A1 or A2, wherein the prompt acquisition means acquires, in addition to the first prompt and the second prompt, a fourth prompt representing an event that must not be included in the digest; and further comprises fourth similarity acquisition means for acquiring a fourth similarity indicating a degree of similarity between features of the image data and features of the fourth prompt by inputting the image data and the fourth prompt into a calculator that calculates the similarity between features of the image data and features of the prompt; and wherein the score calculation means calculates a score indicating that the image data should not be included in the digest if the fourth similarity is equal to or greater than a predetermined threshold.

[0095] (Supplementary Note A4) The information processing device according to any one of Supplementary Notes A1 to A3, further comprising: a digest generation unit configured to generate the digest including image data for which the score calculated by the score calculation unit satisfies a predetermined condition.

[0096] (Appendix A5) The information processing device according to any one of Appendices A1 to A4, wherein the image data is data obtained by photographing a subject, and the digest is data representing an image used to support decision-making regarding the subject.

[0097] (Supplementary Note A6) The information processing device according to any one of Supplementary Notes A1 to A5, wherein the calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt is a visual language model generated by machine learning.

[0098] (Supplementary Note A7) The information processing device according to any one of Supplementary Notes A1 to A6, wherein the first prompt and the second prompt each include at least one of text and image data.

[0099] (Appendix A8) The information processing device according to any one of Appendices A1 to A7, further comprising: a score output unit configured to output the score calculated by the score calculation unit.

[0100] [Appendix B] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0101] (Appendix B1) An information processing method including: an image acquisition process in which at least one processor acquires image data; a prompt acquisition process in which the at least one processor acquires a first prompt representing an event to be included in the digest and a second prompt representing an event to be excluded from the digest; a first similarity acquisition process in which the at least one processor acquires a first similarity indicating a degree of similarity between a feature of the image data and a feature of the first prompt, the first similarity being obtained by inputting the image data and the first prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt; a second similarity acquisition process in which the at least one processor acquires a second similarity indicating a degree of similarity between a feature of the image data and a feature of the second prompt, the second similarity being obtained by inputting the image data and the second prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt; and a score calculation process in which the at least one processor calculates a score for whether to include the image data in the digest using the first similarity and the second similarity.

[0102] (Appendix B2) The information processing method described in Appendix B1, wherein in the prompt acquisition process, the at least one processor acquires, in addition to the first prompt and the second prompt, a third prompt representing an event that must be included in the digest; the at least one processor further includes a third similarity acquisition process in which the at least one processor acquires a third similarity indicating the degree of similarity between the features of the image data and the features of the third prompt by inputting the image data and the third prompt into a calculator that calculates the similarity between the features of the image data and the features of the prompt; and in the score calculation process, the at least one processor calculates a score indicating that the image data should be included in the digest if the third similarity is equal to or greater than a predetermined threshold.

[0103] (Appendix B3) The information processing method described in Appendix B1 or B2, wherein in the prompt acquisition process, the at least one processor acquires, in addition to the first prompt and the second prompt, a fourth prompt representing an event that must not be included in the digest; the at least one processor further includes a fourth similarity acquisition process in which the at least one processor acquires a fourth similarity indicating the degree of similarity between the features of the image data and the features of the fourth prompt by inputting the image data and the fourth prompt into a calculator that calculates the similarity between the features of the image data and the features of the prompt; and in the score calculation process, the at least one processor calculates a score indicating that the image data should not be included in the digest if the fourth similarity is equal to or greater than a predetermined threshold.

[0104] (Supplementary Note B4) The information processing method according to any one of Supplementary Notes B1 to B3, further comprising a digest generation process in which the at least one processor generates the digest including image data for which the score calculated in the score calculation process satisfies a predetermined condition.

[0105] (Appendix B5) The information processing method according to any one of Appendices B1 to B4, wherein the image data is data obtained by photographing a subject, and the digest is data representing an image used to support decision-making regarding the subject.

[0106] (Supplementary Note B6) The information processing method according to any one of Supplementary Notes B1 to B5, wherein the calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt is a visual language model generated by machine learning.

[0107] (Supplementary Note B7) The information processing method according to any one of Supplementary Notes B1 to B6, wherein the first prompt and the second prompt each include at least one of text and image data.

[0108] (Supplementary Note B8) The information processing method according to any one of Supplementary Notes B1 to B7, further including a score output process in which the at least one processor outputs the score calculated in the score calculation process.

[0109] [Appendix C] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims. (Appendix C1) An information processing program that causes a computer to function as an information processing device, the information processing program causing the computer to function as: an image acquisition means that acquires image data; a prompt acquisition means that acquires a first prompt representing an event that is requested to be included in the digest and a second prompt representing an event that is requested not to be included in the digest; a first similarity acquisition means that acquires a first similarity indicating a degree of similarity between a feature of the image data and a feature of the first prompt by inputting the image data and the first prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt; a second similarity acquisition means that acquires a second similarity indicating a degree of similarity between a feature of the image data and a feature of the second prompt by inputting the image data and the second prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt; and a score calculation means that calculates a score for whether to include the image data in the digest using the first similarity and the second similarity.

[0110] (Appendix C2) The information processing program described in Appendix C1, wherein the prompt acquisition means acquires, in addition to the first prompt and the second prompt, a third prompt representing an event that must be included in the digest, and the computer further functions as a third similarity acquisition means that acquires a third similarity indicating the degree of similarity between the features of the image data and the features of the third prompt by inputting the image data and the third prompt into a calculator that calculates the similarity between the features of the image data and the features of the prompt, and the score calculation means calculates a score indicating that the image data should be included in the digest if the third similarity is equal to or greater than a predetermined threshold.

[0111] (Appendix C3) The information processing program described in Appendix C1 or C2, wherein the prompt acquisition means acquires, in addition to the first prompt and the second prompt, a fourth prompt representing an event that must not be included in the digest, and the computer further functions as a fourth similarity acquisition means that acquires a fourth similarity indicating the degree of similarity between the features of the image data and the features of the fourth prompt by inputting the image data and the fourth prompt into a calculator that calculates the similarity between the features of the image data and the features of the prompt, and the score calculation means calculates a score indicating that the image data should not be included in the digest if the fourth similarity is equal to or greater than a predetermined threshold.

[0112] (Appendix C4) The information processing program according to any one of Appendices C1 to C3, further causing the computer to function as a digest generation means for generating the digest including image data for which the score calculated by the score calculation means satisfies a predetermined condition.

[0113] (Appendix C5) The information processing program according to any one of appendices C1 to C4, wherein the image data is data obtained by photographing a subject, and the digest is data representing an image used to support decision-making regarding the subject.

[0114] (Supplementary Note C6) The information processing program according to any one of Supplementary Notes C1 to C5, wherein the calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt is a visual language model generated by machine learning.

[0115] (Supplementary Note C7) The information processing program according to any one of Supplementary Notes C1 to C6, wherein the first prompt and the second prompt each include at least one of text and image data.

[0116] (Appendix C8) The information processing program according to any one of Appendices C1 to C7, which causes the computer to further function as score output means for outputting the score calculated by the score calculation means.

[0117] [Appendix D] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims. (Appendix D1) An information processing device comprising at least one processor, the at least one processor executing: an image acquisition process for acquiring image data; a prompt acquisition process for acquiring a first prompt representing an event that is requested to be included in a digest and a second prompt representing an event that is requested not to be included in the digest; a first similarity acquisition process for acquiring a first similarity indicating a degree of similarity between a feature amount of the image data and a feature amount of the first prompt, the first similarity being obtained by inputting the image data and the first prompt into a calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt; a second similarity acquisition process for acquiring a second similarity indicating a degree of similarity between a feature amount of the image data and a feature amount of the second prompt, the second similarity being obtained by inputting the image data and the second prompt into a calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt; and a score calculation process for calculating a score for whether to include the image data in the digest using the first similarity and the second similarity.

[0118] The information processing device may further include a memory, and the memory may store a program for causing the at least one processor to execute each of the processes.

[0119] (Appendix D2) The information processing device described in Appendix D1, wherein in the prompt acquisition process, the at least one processor acquires, in addition to the first prompt and the second prompt, a third prompt representing an event that must be included in the digest; the at least one processor further executes a third similarity acquisition process to acquire a third similarity indicating the degree of similarity between the features of the image data and the features of the third prompt by inputting the image data and the third prompt into a calculator that calculates the similarity between the features of the image data and the features of the prompt; and in the score calculation process, the at least one processor calculates a score indicating that the image data should be included in the digest if the third similarity is equal to or greater than a predetermined threshold.

[0120] (Appendix D3) The information processing device described in Appendix D1 or D2, wherein in the prompt acquisition process, the at least one processor acquires, in addition to the first prompt and the second prompt, a fourth prompt representing an event that must not be included in the digest; the at least one processor further executes a fourth similarity acquisition process to acquire a fourth similarity indicating the degree of similarity between the features of the image data and the features of the fourth prompt by inputting the image data and the fourth prompt into a calculator that calculates the similarity between the features of the image data and the features of the prompt; and in the score calculation process, the at least one processor calculates a score indicating that the image data should not be included in the digest if the fourth similarity is equal to or greater than a predetermined threshold.

[0121] (Supplementary Note D4) The information processing device according to any one of Supplementary Notes D1 to D3, wherein the at least one processor further executes a digest generation process to generate the digest including image data for which the score calculated in the score calculation process satisfies a predetermined condition.

[0122] (Appendix D5) The information processing device according to any one of appendices D1 to D4, wherein the image data is data obtained by photographing a subject, and the digest is data representing an image used to support decision-making regarding the subject.

[0123] (Supplementary Note D6) The information processing device according to any one of Supplementary Notes D1 to D5, wherein the calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt is a visual language model generated by machine learning.

[0124] (Appendix D7) The information processing device according to any one of appendices D1 to D6, wherein the first prompt and the second prompt each include at least one of text and image data.

[0125] (Appendix D8) The information processing device according to any one of Appendices D1 to D7, wherein the at least one processor further executes a score output process that outputs the score calculated in the score calculation process.

[0126] [Appendix E] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims. (Appendix E1) A non-transitory recording medium having recorded thereon an information processing program that causes a computer to function as an information processing device, the non-transitory recording medium having recorded thereon an information processing program that causes the computer to execute: an image acquisition process that acquires image data; a prompt acquisition process that acquires a first prompt representing an event that is requested to be included in the digest and a second prompt representing an event that is requested not to be included in the digest; a first similarity acquisition process that acquires a first similarity indicating a degree of similarity between a feature of the image data and a feature of the first prompt, the first similarity being obtained by inputting the image data and the first prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt; a second similarity acquisition process that acquires a second similarity indicating a degree of similarity between a feature of the image data and a feature of the second prompt, the second similarity being obtained by inputting the image data and the second prompt into a calculator that calculates the similarity between the feature of the image data and the feature of the prompt; and a score calculation process that calculates a score for whether to include the image data in the digest using the first similarity and the second similarity.

[0127] REFERENCE SIGNS 1, 1A, 1B INFORMATION PROCESSING DEVICE 11 IMAGE ACQUISITION UNIT 12 PROMPT ACQUISITION UNIT 13 FIRST SIMILARITY ACQUISITION UNIT 14 SECOND SIMILARITY ACQUISITION UNIT 15 SCORE CALCULATION UNIT 11A IMAGE INPUT UNIT 12A POSITIVE PROMPT INPUT UNIT 13A NEGATIVE PROMPT INPUT UNIT 14A POSITIVE SCORE CALCULATION UNIT 15A NEGATIVE SCORE CALCULATION UNIT 16A, 16B SUMMARY SCORE CALCULATION UNIT 17A SCORE OUTPUT UNIT 18A DIGEST GENERATION UNIT 20A, 20B MEMORY UNIT 22B MUST PROMPT INPUT UNIT 23B NEVER PROMPT INPUT UNIT 24B MUST SCORE CALCULATION UNIT 25B NEVER SCORE CALCULATION UNIT S1, S1A, S1B INFORMATION PROCESSING METHOD S11 IMAGE ACQUISITION PROCESS S12 PROMPT ACQUISITION PROCESS S13 FIRST SIMILARITY ACQUISITION PROCESS S14 SECOND SIMILARITY ACQUISITION PROCESS S15 SCORE CALCULATION PROCESS

Claims

1. An information processing apparatus comprising: an image acquisition unit that acquires image data; a prompt acquisition unit that acquires a first prompt representing an event to be included in a digest and a second prompt representing an event not to be included in the digest; a first similarity acquisition unit that inputs the image data and the first prompt to a calculator that calculates a similarity between a feature amount of the image data and a feature amount of the prompt, and acquires a first similarity indicating a degree of similarity between the feature amount of the image data and the feature amount of the first prompt; a second similarity acquisition unit that inputs the image data and the second prompt to a calculator that calculates a similarity between a feature amount of the image data and a feature amount of the prompt, and acquires a second similarity indicating a degree of similarity between the feature amount of the image data and the feature amount of the second prompt; and a score calculation unit that calculates a score regarding whether to include the image data in the digest using the first similarity and the second similarity.

2. The information processing apparatus according to claim 1, wherein the prompt acquisition unit further acquires a third prompt representing an event that is essential to be included in the digest, in addition to the first prompt and the second prompt, and a third similarity acquisition unit that inputs the image data and the third prompt to a calculator that calculates a similarity between a feature amount of the image data and a feature amount of the prompt, and acquires a third similarity indicating a degree of similarity between the feature amount of the image data and the feature amount of the third prompt; and the score calculation unit calculates a score indicating that the image data is to be included in the digest when the third similarity is equal to or greater than a predetermined threshold.

3. The prompt acquisition means further acquires a fourth prompt representing an event that must not be included in the digest, in addition to the first prompt and the second prompt, and inputs the image data and the fourth prompt to a calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt, and acquires a fourth similarity indicating the degree of similarity between the feature amount of the image data and the feature amount of the fourth prompt. The score calculation means calculates a score indicating that the image data is not included in the digest when the fourth similarity is equal to or greater than a predetermined threshold. The information processing apparatus according to claim 1 or 2.

4. The information processing apparatus according to any one of claims 1 to 3, further comprising a digest generation means for generating a digest including the image data for which the score calculated by the score calculation means satisfies a predetermined condition.

5. The image data is data obtained by photographing a target person, and the digest is data representing an image used for supporting a decision-making regarding the target person. The information processing apparatus according to any one of claims 1 to 4.

6. The calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt is a vision-language model generated by machine learning. The information processing apparatus according to any one of claims 1 to 5.

7. Each of the first prompt and the second prompt includes at least one of text and image data. The information processing apparatus according to any one of claims 1 to 6.

8. The information processing apparatus according to any one of claims 1 to 7, further comprising a score output means for outputting the score calculated by the score calculation means.

9. At least one processor performs an image acquisition process for acquiring image data, a prompt acquisition process for acquiring a first prompt representing an event to be included in the digest and a second prompt representing an event not to be included in the digest, a first similarity acquisition process for acquiring a first similarity indicating the degree of similarity between the feature amount of the image data and the feature amount of the first prompt, which is obtained by inputting the image data and the first prompt to a calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt, a second similarity acquisition process for acquiring a second similarity indicating the degree of similarity between the feature amount of the image data and the feature amount of the second prompt, which is obtained by inputting the image data and the second prompt to a calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt, and a score calculation process for calculating a score as to whether to include the image data in the digest using the first similarity and the second similarity. An information processing method comprising the above steps.

10. A program for causing a computer to function as an information processing apparatus, the program causing the computer to include: image acquisition means for acquiring image data; prompt acquisition means for acquiring a first prompt representing an event to be included in the digest and a second prompt representing an event not to be included in the digest; first similarity acquisition means for acquiring a first similarity indicating the degree of similarity between the feature amount of the image data and the feature amount of the first prompt, the first similarity being obtained by inputting the image data and the first prompt to a calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt; second similarity acquisition means for acquiring a second similarity indicating the degree of similarity between the feature amount of the image data and the feature amount of the second prompt, the second similarity being obtained by inputting the image data and the second prompt to a calculator that calculates the similarity between the feature amount of the image data and the feature amount of the prompt; score calculation means for calculating a score as to whether to include the image data in the digest using the first similarity and the second similarity; an information processing program for causing the computer to function as described above.

Citation Information

Patent Citations

  • Video control group construction method, device, model training method, device and equipment, and video scoring method, device and equipment

    CN116132752A

  • Video content summarization

    JP2020516107A

  • Systems and methods for analysis of surgical videos

    JP2022520701A

  • Summary video generating device and program thereof

    JP2023122672A

  • Method and Apparatus for Video Digest Generation

    US20090022472A1