Information processing device, information processing method, and recording medium
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2026-01-19
- Publication Date
- 2026-07-30
Smart Images

Figure JP2026001423_30072026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Information Processing Method, and Recording Medium
[0001] The present invention relates to an information processing apparatus, an information processing method, and a recording medium.
[0002] For example, Patent Document 1 discloses a command device for smoothly responding to an emergency notification. The command device described in Patent Document 1 includes a notification reception unit, an information reception unit, and a location extraction unit.
[0003] The notification reception unit receives an emergency notification. The information reception unit receives information on a structure obtained at the notification location from the terminal device that transmitted the emergency notification. The location extraction unit extracts, as candidates for the notification location, locations where information on structures in the target area can be obtained, using the structure features indicated by the structure information and a structure database.
[0004] According to the description of Patent Document 1, the location extraction unit may receive information explaining the notification location from the information reception unit. When receiving voice data as information explaining the notification location, the location extraction unit uses an existing voice recognition technology to generate text data representing the voice data-ized voice from the voice data. The location extraction unit extracts information on the notification location from the generated text data using an existing language recognition technology.
[0005] Japanese Patent Application Laid-Open No. 2011-215767
[0006] However, when the police or the like receive a voice notification (notification voice) regarding a crime such as so-called robbery or assault, in order to appropriately handle the crime, not only the notification location but also information on the suspect, such as the current location of the suspect of the crime, is important. Therefore, there is a possibility that the technique described in Patent Document 1 may not be able to appropriately respond to an event such as a crime.
[0007] One of the problems of the present disclosure is to enable appropriate handling of a notification voice.
[0008] An information processing device in one aspect of this disclosure includes: interpretation means for obtaining search conditions for searching for a search target related to a notification voice based on the notification voice; search means for obtaining the result of searching for the search target from captured images based on the search conditions; and output means for outputting target information relating to the search target based on the search result.
[0009] An information processing method in one aspect of this disclosure involves at least one computer acquiring search conditions for searching for a search target related to a notification voice based on the notification voice, acquiring the results of searching for the search target from captured images based on the search conditions, and outputting target information related to the search target based on the search results.
[0010] One aspect of this disclosure is a program that causes at least one computer to perform the following actions: obtain search conditions for searching for a target related to a notification voice based on the notification voice; obtain the results of searching for the target from captured images based on the search conditions; and output target information related to the target based on the search results.
[0011] According to one example of this disclosure, it becomes possible to appropriately address the audio notification.
[0012] This is a block diagram showing an example configuration of the first information processing device according to this disclosure. This is a flowchart showing an example of the processing operation of the first information processing device according to this disclosure. This is a block diagram showing an example configuration of the first information processing system according to this disclosure. This is a block diagram showing an example configuration of the first interpretation unit according to this disclosure. This is a diagram showing an example configuration when the speech recognition model is provided in an external information processing device. This is a diagram showing an example of the first sentence. This is a diagram showing an example configuration when the first large-scale language model is provided in an external information processing device. This is a diagram showing an example of a prompt and a search condition. This is a flowchart showing a detailed example of the search condition acquisition process according to this disclosure. This is a block diagram showing an example configuration of the first search unit according to this disclosure. This is a diagram showing an example configuration when the first large-scale visual language model is provided in an external information processing device. This is a diagram showing a first example of a prompt and a search result. This is a block diagram showing an example of the physical configuration of the first information processing device according to this disclosure. This is a block diagram showing an example configuration of the second information processing system according to this disclosure. This is a diagram showing an example of second sentence information. This is a block diagram showing an example configuration of the second information processing device according to this disclosure. This is a block diagram showing an example configuration of the second search unit according to this disclosure. This is a diagram showing an example configuration when the second large-scale language model is provided in an external information processing device. This is a diagram showing a second example of a prompt and a search result. This is a block diagram showing an example configuration of the third information processing device according to this disclosure. This is a block diagram showing an example configuration of the second interpreter according to this disclosure. This is a diagram showing an example configuration when the third large-scale language model is provided on an external information processing device. This is a diagram showing an example of a prompt and a summary. This is a flowchart showing an example of the processing operation of the second information processing device according to this disclosure.
[0013] The embodiments relating to this disclosure will be described below with reference to the drawings. In this disclosure, the drawings are associated with one or more embodiments. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted as appropriate.
[0014] <Embodiment 1> As shown in Figure 1, the information processing device 100 according to the present disclosure comprises an interpretation unit 110, a search unit 120, and an output unit 130.
[0015] The interpretation unit 110 obtains search conditions for searching for a target related to the reported audio based on the reported audio. The search unit 120 obtains the results of searching for the target from the captured images based on the search conditions. The output unit 130 outputs target information related to the search target based on the search results.
[0016] The information processing device 100 performs information processing as shown in the flowchart in Figure 2.
[0017] The interpretation unit 110 obtains search conditions for searching for a target related to the notification voice based on the notification voice (step S110). The search unit 120 obtains the results of searching for the target from the captured image based on the search conditions (step S120). The output unit 130 outputs target information related to the target based on the search results (step S130).
[0018] This information processing device 100 can search for a target from captured images based on a voice notification and output information suitable for responding to the notification. Therefore, it becomes possible to respond appropriately to the notification.
[0019] The following describes a detailed example of the information processing device 100.
[0020] (Detailed example) The information processing device 100 is a device that outputs information for dealing with a notification. The information processing device 100 may be provided in the information processing system S1, for example, as shown in Figure 3. The information processing system S1 includes a plurality of imaging devices CM, a plurality of personnel terminals TD, an image storage device DB1, and the information processing device 100.
[0021] The information processing device 100, the image storage device DB1, the imaging device CM, and the personnel terminal TD are all connected to each other via the communication network NT. These devices send and receive information from each other via the communication network NT. The communication network NT is a communication network configured as wired, wireless, or a combination thereof.
[0022] Note that there may be only one imaging device CM. Also, there may be only one personnel terminal TD. The image storage device DB1 may be provided in the information processing device 100. The image storage device DB1 may consist of multiple storage devices.
[0023] (Regarding the reporting audio) The reporting audio is the audio related to the reporting. For example, the reporting audio includes at least one of the following: the audio spoken by the person making the report and the audio of the person receiving the report speaking to the person making the report.
[0024] Reports concern incidents such as crimes, accidents, and lost children. Crimes include, for example, snatching, assault, and robbery. Accidents include, for example, hit-and-run accidents, collisions between vehicles, and collisions between vehicles. In such cases, the person who receives the report is, for example, a police officer.
[0025] Please note that the incidents related to reporting are not limited to those exemplified here. Also, the person receiving the report is not limited to police officers or other police personnel; for example, it could be an employee of a designated organization such as a security company.
[0026] Reporting is done, for example, by the reporter using the reporter terminal RT. The reporter terminal RT is a device held by the reporter.
[0027] Here, the terminal is, for example, a smartphone. The terminal may also be a tablet, a personal computer, a telephone, etc. The same applies to the terminal below.
[0028] The caller terminal RT is equipped with a call function for making voice calls. The caller terminal RT may also be equipped with a location tracking function that determines the current location of the caller terminal RT using GPS, wireless LAN (Local Area Network), etc. The call may include the voice content and the current location of the caller terminal RT. The call may further include terminal identification information such as the caller terminal RT's telephone number.
[0029] (Regarding the shooting device CM) Each shooting device CM is a camera or the like that acquires captured images. Each shooting device CM acquires video (moving images), still images, etc., of the shooting area as captured images. Still images may be individual frame images that make up a moving image, frame images extracted from a moving image at predetermined time intervals, etc.
[0030] Each of the recording devices CM may be a security camera installed in an urban area, building, etc. Each of the recording devices CM may be fixed or movable.
[0031] Each of the imaging devices CM outputs imaging information, including, for example, the captured image.
[0032] The shooting information may further include shooting-related information associated with the captured image. Shooting-related information is information associated with the captured image and may include at least one of the following: for example, the shooting device ID (identification), shooting time, shooting position, shooting direction, resolution, and shooting range. If the captured image is a moving image, the shooting-related information may be linked to the moving image in such a way that the content of the shooting-related information for each frame image constituting the moving image can be identified.
[0033] The imaging device ID is information used to identify the imaging device CM that captured the image.
[0034] The shooting date is the time when the image was taken.
[0035] The period includes, for example, the time expressed using year, month, day, hour, and minute. However, the method of expressing the period is not limited to this. The same applies to the following:
[0036] The shooting position is the position of the imaging device CM. The shooting position may also be the position of the shooting area captured by the imaging device CM.
[0037] A location includes at least one of the following: latitude and longitude, height (e.g., elevation), address, etc. The method of representing a location is not limited to these. The same applies to the following:
[0038] The shooting direction is the direction in which the shooting device CM captured the image.
[0039] The direction includes at least one of, for example, a vector corresponding to latitude, longitude, and altitude (e.g., elevation), an azimuth corresponding to north, south, east, and west, an aspect that is the direction towards a city, a facility, etc., a clock position, an azimuth angle, etc. The method of representing the direction is not limited to this. The same applies hereinafter.
[0040] The resolution is the resolution of the captured image, and is represented by, for example, the total number of pixels, the number of pixels in each of the vertical and horizontal directions, etc. The resolution is determined according to, for example, the settings, specifications, etc. of the imaging device CM that captured the captured image.
[0041] The imaging range is the range that can be imaged by the imaging device CM that captured the captured image, and is, for example, the range that can be imaged on the ground, on the road surface, on the floor surface, etc. The imaging range may be represented in a predetermined shape. This predetermined shape may be a polygon such as a quadrilateral, and in this case, the imaging range may include the coordinates of the vertices of the polygon. The coordinates are represented by, for example, latitude and longitude.
[0042] (Regarding the image storage device DB1) The image storage device DB1 stores the captured images captured by each of the imaging devices CM. The image storage device DB1 may store the current captured image, that is, the real-time captured image, and the past captured images. The image storage device DB1 may further store imaging-related information. That is, the image storage device DB1 may store imaging information.
[0043] (Regarding the personnel terminal TD) The personnel terminal TD is a terminal held by a predetermined person. The predetermined persons may be plural, and each may hold the personnel terminal TD.
[0044] The personnel terminal TD is an output destination of the information (e.g., target information, etc.) output by the information processing device 100. In addition to the communication function of acquiring the information (e.g., target information, etc.) output by the information processing device 100, the personnel terminal TD may be provided with at least one of a call function, a position identification function, a display function, etc. The position identification function of the personnel terminal TD is a function of identifying the current position of the personnel terminal TD using GPS, wireless LAN (Local Area Network), etc.
[0045] (Regarding the information processing apparatus 100) As described above with reference to FIG. 1, the information processing apparatus 100 includes an interpretation unit 110, a search unit 120, and an output unit 130.
[0046] (Regarding the interpretation unit 110) As described above, the interpretation unit 110 acquires search conditions for searching for a search target related to the reported voice based on the reported voice.
[0047] The search target is a person or object related to the reported voice and is a target for search using a captured image. There may be a plurality of search targets.
[0048] When the report is related to a crime, the search target is, for example, a suspect of the crime. When the report is related to an accident, the search target is, for example, at least one of the parties to the accident, a vehicle that caused the accident, etc. When the report is related to a missing person, the search target is, for example, the missing person. Note that the search target is not limited to those exemplified here and may be changed according to the event related to the report.
[0049] (Regarding a detailed example of the interpretation unit 110) As shown in FIG. 4, the interpretation unit 110 may include a first sentence acquisition unit 111 and a search condition acquisition unit 112. The first sentence acquisition unit 111 acquires a first sentence obtained by formulating the content of the reported voice. The search condition acquisition unit 112 acquires search conditions based on the first sentence.
[0050] (Regarding a configuration example of the first sentence acquisition unit 111) The first sentence acquisition unit 111 acquires the reported voice. The first sentence acquisition unit 111 may acquire report accompanying information together with the reported voice. The report accompanying information is information accompanying the report and includes, for example, at least one of the report time, the current position of the reporter terminal RT used for the report, and the terminal identification information.
[0051] The report time is the time when the report was made, for example, the time when the person in charge received the report.
[0052] The first sentence acquisition unit 111 acquires the reported voice using the acquired reported voice and a voice recognition model. The first sentence acquisition unit 111 inputs the reported voice into the voice recognition model to acquire the first sentence.
[0053] A speech recognition model is a machine learning model that has been trained to convert speech into text. General techniques may be used to build a speech recognition model.
[0054] The speech recognition model may be included in the first text acquisition unit 111. The speech recognition model may also be provided in the information processing devices COM1 that are connected to each other via a communication network NT, as shown in Figure 5.
[0055] Figure 6 shows an example of the first sentence. The first sentence shown in the figure is an example of the first sentence obtained from a report made to the police by a witness who saw a snatching incident in front of XX Supermarket in front of Station A. In the figure, "◆:" indicates that the following sentence is spoken by a police officer. In the figure, "◇:" indicates that the following sentence is spoken by the caller.
[0056] (Regarding the search condition acquisition unit 112) The search condition acquisition unit 112 acquires search conditions using the first sentence and the first large-scale language model. The search condition acquisition unit 112 inputs the first sentence into the first large-scale language model and acquires search conditions.
[0057] The search condition acquisition unit 112 may further acquire search conditions using the information associated with the notification. The search condition acquisition unit 112 may also acquire search conditions by inputting the first sentence and the information associated with the notification into the first large-scale language model.
[0058] The first large-scale language model is a large-scale language model used to create search conditions.
[0059] Hereafter, "Large-Scale Language Model" and "First Large-Scale Language Model" will also be referred to as "LLM" and "First LLM," respectively.
[0060] The first LLM may be an LLM built using general-purpose techniques. An LLM is a machine learning model that outputs text data in response to an input prompt. The prompt input to the LLM is text data.
[0061] The first LLM may be included in the search condition acquisition unit 112. The first LLM may be provided in the information processing device COM2, which is connected to each other via the communication network NT, as shown in Figure 7. Note that the information processing device COM2 may be the same device as the information processing device COM1 described above.
[0062] For example, when the first LLM receives a prompt containing the first sentence, it interprets the first sentence and outputs a search condition as a sentence corresponding to the prompt. In this case, the prompt includes, for example, an instruction to create a search condition for searching for a target corresponding to the first sentence.
[0063] For example, when the first LLM receives a prompt that includes additional information accompanying the notification, it interprets the first sentence and the information accompanying the notification together and outputs a search condition as a sentence corresponding to the prompt. In this case, the prompt includes, for example, an instruction to create a search condition for searching for a target corresponding to the first sentence and the information accompanying the notification.
[0064] The prompts entered into the first LLM may be created by the search condition acquisition unit 112 or the user according to a predetermined template.
[0065] The template may include information that identifies the location for setting the first sentence. This allows the search unit 120 to automatically create a prompt to input into the first LLM using a predetermined template and the first sentence acquired by the first sentence acquisition unit 111. The template may further include information that identifies the location for setting accompanying notification information. This allows the search unit 120 to automatically create a prompt to input into the first LLM using the accompanying notification information.
[0066] The user may modify part of the prompt created by the search condition acquisition unit 112.
[0067] (Regarding search conditions) Search conditions are the conditions for searching for the target, and are created based on the notification audio as described above. These search conditions are then used to search for the target from the captured images.
[0068] The search conditions may include (1) image selection conditions and (2) target search conditions.
[0069] (1) The image selection criteria are the conditions for selecting the images to be used in the search for the target of the search. The image selection criteria include, for example, at least one of the search start time, search starting point, and search priority direction.
[0070] The search start time is the time when the search for the target of the search begins.
[0071] The search starting point is the location from which the search for the target object begins.
[0072] The search priority direction is the direction in which the search range for the target object is changed. Changing the search range includes, for example, at least one of expansion and movement.
[0073] (2) The target search conditions are the conditions that the search target must satisfy. The target search conditions include, for example, the type of search target and at least one of the attributes of the search target.
[0074] The search target type is the type of object being searched for. Examples of search target types include people, cars, bicycles, etc.
[0075] The target attribute is the attribute of the object being searched. The target attribute may differ depending on the type of object being searched.
[0076] When the search target type is "person," the search target attributes include, for example, at least one of the following: age group, gender, body type, height, clothing, etc. Age group includes, for example, at least one of the lower age limit and upper age limit. Height includes, for example, at least one of the lower height limit and upper height limit. Clothing includes, for example, at least one of the upper body clothing color, upper body clothing type, lower body clothing color, and lower body clothing type.
[0077] When the search target type is "automobile," the search target attributes include at least one of the following: color, car name, vehicle type, and license plate number information. The vehicle type may include at least one of the following: regular car, large vehicle, light vehicle, etc. The vehicle type may be defined in a hierarchical structure, for example, including a major category and a minor category. The major category may include at least one of the following: regular car, large vehicle, light vehicle, etc. For example, the minor category for regular car may include at least one of the following: sedan, van, wagon, truck, etc. License plate number information is information contained in the license plate attached to the automobile.
[0078] Figure 8 shows an example of a prompt and search condition.
[0079] The prompt in the figure is an example that includes instructions for creating search conditions to search for a target corresponding to the first sentence and accompanying information of the notification, as illustrated in Figure 6. The accompanying information of the notification includes the notification time, "hh:mm". The type of target to search is "person".
[0080] More specifically, the prompts shown in the diagram include "standard prompts," which are fixed instructions, and "individual prompts," which are different instructions for each report.
[0081] A standard prompt includes instructions regarding processing content, output format, etc. The diagram shows an example where the output format is specified as JSON. However, the output format is not limited to this; for example, it could be text format, etc.
[0082] Each individual prompt includes the first sentence and any accompanying information. While details are omitted in the diagram for clarity, the first sentence in the diagram is set to the text data of the first sentence exemplified in Figure 6.
[0083] The search conditions shown in the figure are an example obtained by inputting the above prompt into the first LLM, and include publicly available information obtainable via the internet, etc. The publicly available information is general information that is publicly available and obtainable via the internet, etc.
[0084] Here, the location of the XX supermarket mentioned in the first sentence is assumed to be at latitude p.p degrees North and longitude q.q degrees East, and its address is xx, n-chome, A-cho. Furthermore, A Station is assumed to be located west of the XX supermarket. This information is publicly available and can be obtained via the internet, etc.
[0085] The first LLM can obtain this publicly available information through the internet or other means, and use it to interpret the first document and the accompanying information of the notification to create search conditions.
[0086] In the figure, the search start time included in the search conditions (image selection conditions) is an example of a search start time expressed as a time, and the notification time is set. The search starting point "p.pN, q.qE" and the search priority direction "East" included in the search conditions (image selection conditions) are set based on the first document and the publicly available information as described above.
[0087] In the diagram, the search conditions (target search conditions) are set based on the first sentence, with conditions corresponding to each of the target attributes of the person being searched. Either or both of the information accompanying the notification, the publicly available information, or both may be used to set the search conditions (target search conditions).
[0088] Note that the search conditions are not limited to the examples given above.
[0089] For example, the start time of the search may differ from the time the report was made. For instance, if the first sentence includes a statement that "a snatching incident occurred about five minutes ago," the start time of the search may be set to a time five minutes or more before the time the report was made.
[0090] For example, the search conditions may include a search range such as radius R [m] in addition to the image selection conditions. The search range may vary depending on the manner in which the subject of the search left the scene, the time elapsed since the incident occurred or was reported, and the conditions at the scene. The manner of movement may include walking, running, or driving away in a car. The conditions at the scene may include the amount of pedestrian traffic and the width of the road.
[0091] (Detailed example of the processing operation of the interpretation unit 110) The interpretation unit 110 may perform a search condition acquisition process (step S110) as shown in Figure 9. The search condition acquisition process (step S110) is started, for example, when the first text acquisition unit 111 acquires the notification voice. Here, notification voice information may be acquired along with the notification voice.
[0092] However, the trigger for starting the search condition acquisition process (step S110) is not limited to this.
[0093] The first text acquisition unit 111 acquires the first text, which is a written representation of the content of the notification audio (step S111).
[0094] For example, the first text acquisition unit 111 uses a speech recognition model to acquire the first text.
[0095] If the first text acquisition unit 111 includes a speech recognition model, the first text acquisition unit 111 inputs the notification voice into the speech recognition model and acquires the first text.
[0096] When a speech recognition model is provided in the information processing device COM1, the first text acquisition unit 111 inputs the notification voice to the speech recognition model by transmitting the notification voice to the information processing device COM1. As a result, the speech recognition model outputs the first text. The first text acquisition unit 111 acquires the first text from the information processing device COM1.
[0097] The search condition acquisition unit 112 acquires search conditions based on the first sentence (step S112). Here, the search condition acquisition unit 112 may acquire search conditions based on the first sentence and the notification voice information.
[0098] For example, as described above, the search condition acquisition unit 112 uses the first LLM to acquire the search conditions.
[0099] If the search condition acquisition unit 112 includes a first LLM, the search condition acquisition unit 112 acquires the search conditions by inputting a prompt containing a first sentence into the first LLM. If notification voice information is acquired, the notification voice information may also be input into the first LLM.
[0100] If the first LLM is provided in the information processing device COM2, the search condition acquisition unit 112 inputs the first sentence into the first LLM by, for example, sending a prompt containing the first sentence to the information processing device COM2. If notification voice information is acquired, a prompt further containing the notification voice information may be sent to the information processing device COM2 and input into the first LLM. As a result, the first LLM outputs the search conditions. The search condition acquisition unit 112 acquires the search conditions from the information processing device COM2.
[0101] (Regarding the search unit 120) As described above, the search unit 120 obtains the results of searching for a target from the captured image based on the search conditions. For example, the search unit 120 obtains the results of searching for a target from the captured image using the search conditions and the first large-scale visual language model.
[0102] The first large-scale visual language model is a large-scale visual language model used for searching for targets.
[0103] Hereafter, the "Large-Scale Visual Language Model" and the "First Large-Scale Visual Language Model" will also be referred to as "VLM" and "First VLM," respectively. VLM is an abbreviation for Vision Language Model.
[0104] The first VLM may be a VLM constructed using general techniques. The VLM is a machine learning model that integrates the functions of Artificial Intelligence (AI) and LLM. The VLM outputs data in response to input prompts. The prompts input to the VLM may include not only text data but also images. The data output from the VLM includes at least one of text and images.
[0105] (Details of the search unit 120) The search unit 120 includes, for example, an image selection unit 121 and a search result acquisition unit 122, as shown in Figure 10. The image selection unit 121 selects images from the captured images to be used for searching for the target. The search result acquisition unit 122 acquires the results of searching for the target from the selected captured images.
[0106] (Regarding the configuration example of the image selection unit 121) The image selection unit 121 uses image selection conditions to select images from among the captured images to be used for searching for the target of search.
[0107] For example, the image selection unit 121 selects captured images that satisfy the search conditions. The image selection unit 121 may also select captured images that satisfy the search conditions, prioritizing them in order of their shooting location being closest to the search starting point. The selected results may be in the form of a list in which the captured images are arranged in descending order of priority.
[0108] Images that satisfy the search conditions are, for example, images that satisfy (1) spatial conditions and (2) temporal conditions.
[0109] (1) The spatial condition is, for example, that the shooting location is included in the search range. The search range may be a predetermined range (spatial range) from the search starting point of the search condition.
[0110] The search range may include a range predetermined from the search starting point of the search conditions, and a range obtained by moving within that range in the search priority direction.
[0111] Here, the shooting position may be the position of the shooting device CM, or it may be the position of the shooting area that the shooting device CM is shooting. If the shooting position is the position of the shooting area, the fact that the shooting position is included in a predetermined range from the search starting point of the search conditions may mean that at least a part of the shooting position is included in a predetermined range from the search starting point of the search conditions. The position of the shooting area may be included in the shooting-related information as the shooting position, or it may be calculated by the image selection unit 121 using the position of the shooting device CM, the shooting direction, the shooting range, etc.
[0112] (2) The temporal condition is, for example, that the time of shooting falls within a predetermined range (temporal range) from the time the search begins.
[0113] (Example of the configuration of the search result acquisition unit 122) The search result acquisition unit 122 uses the target search conditions and the first VLM to acquire the results of searching for a target from the selected captured images.
[0114] For example, the search result acquisition unit 122 creates a prompt to be input to the first VLM. The prompt to be input to the first VLM includes target search conditions and image identification information. The prompt to be input to the first VLM may also include, for example, an instruction to search for a search target that matches the search conditions from among the captured images.
[0115] Image identification information is information used to identify the captured images used in the search for the target object, for example, the captured images used in the search for the target object. The image identification information is set in order of selected captured images, for example, according to priority.
[0116] Image identification information is not limited to the captured image used to search for the target, but may also include, for example, the location of the captured image. The location of the captured image is the address of the memory unit where the captured image is stored. The address may include the address in the communication network NT (network address), or it may include the directory name, file name, etc.
[0117] The search result acquisition unit 122 may create prompts to be input to the first VLM according to a predetermined template. The template may include information that identifies the location where each condition included in the target search conditions (see Figure 8) is set. As a result, the search result acquisition unit 122 may automatically create prompts to be input to the first VLM using the predetermined template and the target search conditions acquired by the search condition acquisition unit 112. The user may also create prompts or modify parts of the prompts created by the search result acquisition unit 122.
[0118] For example, the search result acquisition unit 122 inputs the created prompt to the first VLM. Upon receiving the prompt, the first VLM comprehensively interprets the target search conditions and the captured image to search for a target that matches the search conditions within the captured image. Further interpretation of the target search conditions and the captured image may include information associated with the image capture, publicly available information obtainable via the internet, etc. The first VLM outputs the search results.
[0119] The first VLM may be included in the search result acquisition unit 122. The first VLM may be provided in the information processing device COM3, which is connected to the search result acquisition unit 122 via a communication network NT, as shown in Figure 11. The information processing device COM3 may be the same device as one or both of the information processing devices COM1 and COM2 described above.
[0120] For example, the search result acquisition unit 122 inputs a prompt to the first VLM and acquires the result of searching for the target from the captured image.
[0121] In detail, if the search result acquisition unit 122 includes the first VLM, the search result acquisition unit 122 acquires the results of searching for the target from the captured image by inputting a prompt to the first VLM.
[0122] When the first VLM is provided in the information processing device COM3, the search unit 120 inputs a prompt to the first VLM by sending a prompt to the information processing device COM3. As a result, the first VLM outputs the result of searching for the target from the captured image. The search result acquisition unit 122 acquires the search result from the information processing device COM3.
[0123] The search result is information indicating that the search target could not be found. If the search target could not be found, the search result acquisition unit 122 may recreate the prompt. The image identification information of the recreated prompt includes the next highest priority image after the image used in the previous search. As described above, the selected images are selected in order of priority, starting with those closest to the search starting point. Therefore, the search range can be changed by changing the images used to search for the search target according to their priority.
[0124] Furthermore, once the search using all selected images is complete, the image selection unit 121 may re-select images in order of priority, starting with those closest to the search starting point, in order to expand the search range. In the re-selection of images, the search range for spatial conditions may be broader than that used in the previous selection of images. The image identified by the image identification information may be the image with the highest priority among the re-selected images that have not yet been used in the search for the target of the search.
[0125] The recreated prompt may be the same as the previous prompt, except for the image identification information. The target search conditions may be relaxed in the recreated prompt.
[0126] The relaxed target search conditions may also be the target attributes being searched. In this case, the recreated prompt may include target attributes that have been modified to search using new conditions that encompass the conditions of the target attributes used in the previous search. Which target attributes to modify and how may be predetermined. This relaxes the target search conditions so that, if the target cannot be found, the search uses new conditions that encompass the conditions of the target attributes used in the previous search.
[0127] The search result acquisition unit 122 inputs the recreated prompt to the first VLM and acquires the search results. The search result acquisition unit 122 repeats the process of recreating the prompt and acquiring the search results using the first VLM until, for example, a termination condition is met. The termination condition may include at least one of the following: the search target is found; the search is completed using all of the selected captured images; or the search target has been searched using a predetermined number of captured images.
[0128] If the target of the search is found, the search results may include at least one of the following: the location of the target of the search, the direction of movement of the target of the search, the time when the image in which the target of the search was found (i.e., the image in which the target of the search was taken), the image of the target of the search, the image in which the event related to the report was found, and the location and time when the event related to the report occurred.
[0129] An image in which a search target is found is at least one image containing the search target. An image in which an event related to a report is found is at least one image containing the search target at the time the event related to the report occurred.
[0130] The search results may include a map image. The map image is an image that shows at least one of location and direction, such as a map of a map or a map of a building. The location included in the map image is at least one of the following: for example, the location of the object being searched for, the location where the event occurred, etc. The direction is, for example, the direction of movement of the object being searched for.
[0131] Note that the information included in the search results is not limited to what is exemplified here.
[0132] Figure 12 shows a first example of a prompt and search results.
[0133] The prompts in the figure are examples of instructions that include commands to search for a target according to the search conditions exemplified in Figure 8, from the captured image included in the image identification information.
[0134] The prompts shown in the figure include "standard prompts," which are fixed instructions, "individual prompts," which are instructions that differ for each report, and [input image], which is an instruction for the image to be captured for the search.
[0135] Standard prompts include instructions regarding processing content, output format, etc.
[0136] The diagram shows an example where the output format is specified as a text format intended for wireless transmission.
[0137] Although not shown in the diagram, the standard prompt may also include instructions such as, "If not found, increase the upper age limit by 10 years and decrease the lower age limit by 10 years." This allows for a relaxation of the age-related conditions for the search target. This instruction is an example of relaxing the search conditions so that, if the search target cannot be found, the search is performed using new conditions that encompass the search target attribute conditions used in the previous search.
[0138] Each individual prompt has target search conditions set. The input image has the captured image used for searching for the target set as image identification information.
[0139] The search results shown in the figure are an example of information output from the first VLM when the search target is found after inputting the prompt shown in the figure to the first VLM. The figure shows an example where the search results are output as text according to the instructions of the prompt.
[0140] (Regarding the output unit 130) As described above, the output unit 130 outputs target information related to the search target based on the search results.
[0141] The output unit 130 outputs target information about the target based on the search results when it finds a target among the captured images.
[0142] The target information is information relating to the search target and includes, for example, at least one of the following (A) to (J): (A) Location of the search target (B) Direction of movement of the search target (C) Time when the image in which the search target was found was taken (D) Image of the search target (E) At least part of the search conditions (F) Event related to the report (G) Location where the event related to the report occurred (H) Time when the event related to the report occurred (I) Image in which the event related to the report was found (J) Report audio
[0143] The image in which the target of search is found is an image that contains the target of search. The image of the target of search may be the image in which the target of search was found, or it may be an image extracted from this image containing the region of the target of search.
[0144] The target information may include map images showing the location of the object to be searched, such as maps or building layouts.
[0145] The output destinations to which the output unit 130 outputs the target information are the display unit (not shown) of the information processing device 100, other devices, etc. The other devices may be personnel terminals TD.
[0146] The output unit 130 may transmit the information to personnel terminals TD held by personnel within a predetermined range from the location of the search target. In other words, the target information may be transmitted to one or more personnel terminals TD held by one or more personnel within a predetermined range from the location of the search target.
[0147] (Example of physical configuration of the information processing device 100) Figure 13 shows an example of the physical configuration of the information processing device 100. Physically, the information processing device 100 includes, for example, a bus 1010, a processor 1020, a memory 1030, a storage device 1040, a network interface 1050, an input interface 1060, and an output interface 1070.
[0148] Bus 1010 is a data transmission path for the processor 1020, memory 1030, storage device 1040, network interface 1050, input interface 1060, and output interface 1070 to send and receive data to and from each other. However, the method of connecting the processor 1020 and the other components to each other is not limited to bus connection.
[0149] Processor 1020 is a processor implemented using components such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit).
[0150] Memory 1030 is a main memory device implemented as RAM (Random Access Memory), etc.
[0151] The storage device 1040 is an auxiliary storage device implemented as an HDD (Hard Disk Drive), SSD (Solid State Drive), memory card, or ROM (Read Only Memory). The storage device 1040 stores program modules for realizing the functions of the device equipped with it. The processor 1020 reads each of these program modules into the memory 1030 and executes them, thereby realizing the functions corresponding to those program modules.
[0152] The network interface 1050 is an interface for connecting a device equipped with it to a communication network.
[0153] The input interface 1060 is an interface for the user to input information. The input interface 1060 consists of, for example, a touch panel, a keyboard, a mouse, and the like.
[0154] The output interface 1070 is an interface for presenting information to the user. The output interface 1070 is composed of, for example, a liquid crystal panel, an organic EL (Electro-Luminescence) panel, and the like.
[0155] Embodiment 1 has been described above.
[0156] (Function and Effects) According to this embodiment, the information processing device 100 comprises an interpretation unit 110, a search unit 120, and an output unit 130.
[0157] The interpretation unit 110 obtains search conditions for searching for a target related to the reported audio based on the reported audio. The search unit 120 obtains the results of searching for the target from the captured images based on the search conditions. The output unit 130 outputs target information related to the search target based on the search results.
[0158] This allows the system to search for the target of the investigation within the captured images based on the voice notification, and output information suitable for responding to the notification. Therefore, it becomes possible to respond appropriately to the notification.
[0159] According to this embodiment, the target information includes at least one of the following: the location of the target to be searched, the direction of movement of the target to be searched, the time when the captured image including the target to be searched was taken, the image of the target to be searched, at least a part of the search conditions, the event related to the notification, the location where the event related to the notification occurred, the time when the event related to the notification occurred, the captured image in which the event related to the notification was discovered, and the notification audio.
[0160] This allows us to obtain relevant information that is useful for dealing with reports. Therefore, it becomes possible to respond to reports appropriately.
[0161] According to this embodiment, the captured image used for searching for the target is changed if the target cannot be found.
[0162] This increases the likelihood of finding the target of the search. Consequently, it increases the likelihood of responding appropriately to reports.
[0163] According to this embodiment, the modification of the captured image used for searching for the target of search includes at least one of expanding and moving the search range.
[0164] This increases the likelihood of finding the target of the search. Consequently, it increases the likelihood of responding appropriately to reports.
[0165] According to this embodiment, the captured image used for searching for the target object is changed based on the search conditions.
[0166] This increases the likelihood of finding the target of the search early. Consequently, it becomes even more likely that reports can be dealt with appropriately.
[0167] According to this embodiment, the search conditions are relaxed if the target cannot be found.
[0168] This increases the likelihood of finding the target of the search. Consequently, it increases the likelihood of responding appropriately to reports.
[0169] According to this embodiment, the interpretation unit 110 includes a first text acquisition unit 111 and a search condition acquisition unit 112. The first text acquisition unit 111 acquires a first text which is a written representation of the content of the notification voice. The search condition acquisition unit 112 acquires search conditions based on the first text.
[0170] This allows for the acquisition of search conditions based on voice notifications. Then, using these search conditions, the system can search for the target within the captured images and output information suitable for responding to the notification. Therefore, it becomes possible to increase the likelihood of appropriately responding to notifications.
[0171] In this embodiment, the first text acquisition unit 111 inputs the notification voice into the speech recognition model to acquire the first text. The search condition acquisition unit 112 inputs the first text into the first large-scale language model to acquire the search conditions.
[0172] This allows for the acquisition of search conditions based on voice notifications. Then, using these search conditions, the system can search for the target within the captured images and output information suitable for responding to the notification. Therefore, it becomes possible to increase the likelihood of appropriately responding to notifications.
[0173] According to this embodiment, the search unit 120 inputs the search conditions into a large-scale visual language model and obtains the results of searching for the target from the captured image.
[0174] This allows for the acquisition of search conditions based on voice notifications. Then, using these search conditions, the system can search for the target within the captured images and output information suitable for responding to the notification. Therefore, it becomes possible to increase the likelihood of appropriately responding to notifications.
[0175] According to this embodiment, the target information is transmitted to one or more personnel terminals TD held by one or more personnel within a predetermined range from the location of the search target.
[0176] This allows information about the target to be shared with one or more individuals within a predetermined range from the target of the search, enabling them to respond to reports. Therefore, it becomes possible to respond appropriately to reports.
[0177] <Embodiment 2> Embodiment 1 described an example in which a VLM is used to search for a target from a captured image. Embodiment 2 describes an example in which an LLM (second LLM) is used to search for a target from a captured image.
[0178] In this embodiment, in order to simplify the explanation, descriptions that overlap with other embodiments will be omitted as appropriate.
[0179] The information processing system S2 according to this embodiment includes, for example, as shown in Figure 14, a plurality of imaging devices CM, a plurality of personnel terminals TD, and an image storage device DB1, similar to those in Embodiment 1. The information processing system S2 further includes a second document storage device DB2, an image document conversion device 240, and an information processing device 200.
[0180] The second document storage device DB2, the image document processing device 240, the information processing device 200, the image storage device DB1, the imaging device CM, and the personnel terminal TD are all connected to each other via the communication network NT. These devices send and receive information from each other via the communication network NT.
[0181] The second document storage device DB2 may be provided in the information processing device 200. The image storage device DB1 may consist of multiple storage devices. At least a portion of the image storage device DB1 and the second document storage device DB2 may be the same storage device. The image document conversion device 240 may be provided in the information processing device 200 as an image document conversion unit that realizes this function. The image document conversion device 240 may be the same device as one or more of the information processing devices COM1 to COM3.
[0182] (Regarding the second document storage device DB2) The second document storage device DB2 is a device that stores the second document.
[0183] The second document is a written description of the content of the captured image, for example, a description of the content of the captured image. If the captured image is a moving image, the second document may also include a description of the content of each frame image that makes up the moving image. The second document storage device DB2 may store second document information that includes the second document, which is a written description of the captured image, and information associated with the capture of the image.
[0184] Figure 15 shows an example of second text information. The second text information shown in the figure is an example that includes the second text and shooting-related information. The shooting-related information shown in the figure includes the shooting device ID "CM1" and the shooting date.
[0185] (About the image-to-text device 240) The image-to-text device 240 is a device that creates a second document from a captured image. The image-to-text device 240 creates a second document from a captured image in real time, for example. However, the timing of generating the second document is not limited to this.
[0186] For example, the image-to-text device 240 includes a VLM and uses the VLM to create a second text.
[0187] The image-to-text device 240 continuously acquires, for example, real-time captured images from multiple imaging devices CM or image storage device DB1. The image-to-text device 240 may also acquire accompanying information along with the captured images.
[0188] The VLM interprets the captured image in real time and outputs the second sentence in real time, following prompts that include instructions for creating a second sentence for the captured image. Further information such as information accompanying the capture and publicly available information obtained through the internet may also be used in interpreting the captured image.
[0189] The image-to-text device 240 stores the outputted second text in the second text storage device DB2. This allows the second text, which is a transcription of images from past captured images to the most recent captured image, to be stored in the second text storage device DB2.
[0190] Note that the method for creating the second document is not limited to the method described here. The method for storing the second document in the second document storage device DB2 is not limited to the method described here.
[0191] (Regarding the information processing device 200) As shown in Figure 16, the information processing device 200 includes an interpretation unit 110 and an output unit 130, and a search unit 220, similar to those in Embodiment 1.
[0192] (Regarding the search unit 220) The search unit 220, similar to the search unit 120 in Embodiment 1, acquires the results of searching for a target from the captured image based on the search conditions.
[0193] The search unit 220 obtains the results of searching for a target from the captured images based on a second sentence which is a written description of the contents of the captured image and the search conditions.
[0194] The search unit 220 uses, for example, a second sentence which is a written description of the contents of the captured image, search conditions, and a second LLM to obtain the results of searching for a target from the captured image. The search unit 220 may also use additional information associated with the image capture to obtain the results of searching for a target from the captured image.
[0195] The second LLM is an LLM used to search for the target of the search. The second LLM may be an LLM constructed using general techniques. The second LLM may be the same LLM as the first LLM. The first LLM and the second LLM differ in at least the input prompts.
[0196] (Detailed example of the search unit 220) The search unit 220 includes, for example, an image selection unit 121 and a search result acquisition unit 222, similar to those in Embodiment 1, as shown in Figure 17. The search result acquisition unit 222 acquires the results of searching for a target from among the selected captured images, similar to the search result acquisition unit 122 in Embodiment 1. The search result acquisition unit 222 differs from the search result acquisition unit 122 in that it uses a second sentence, which is a written description of the selected captured images, to search for a target. That is, the search result acquisition unit 222 uses a second sentence, which is a written description of the selected captured images, to acquire the results of searching for a target from among the selected captured images.
[0197] (Detailed example of the search result acquisition unit 222) The search result acquisition unit 222 uses, for example, the second sentence of the selected captured image, the target search conditions, and the second LLM to acquire the results of searching for the search target from the selected captured image.
[0198] For example, the search result acquisition unit 222 creates a prompt to be input to the second LLM. The prompt to be input to the second LLM includes target search conditions and second document identification information. The prompt to be input to the second LLM may also include, for example, an instruction to search for a search target that matches the search conditions from the second document which is a written description of a captured image.
[0199] The second document identification information is information for identifying the second document used in the search for the target, and is, for example, a second document that describes the captured image used in the search for the target. The second document identification information is set in order, for example, according to the priority of the captured images, with the second documents that describe the selected captured images being set in order. If the second document is a long text exceeding a predetermined number of characters, multiple prompts may be created for a single captured image by dividing that second document.
[0200] The second document identification information is not limited to the second document, which is a written description of the captured image used to search for the target of the search, but may also include, for example, the location of the second document. The location of the second document is the address of the memory unit where the second document is stored. The address may include an address in the communication network NT (network address), a directory name, a file name, etc.
[0201] The search result acquisition unit 222 may create prompts to be input to the second LLM according to a predetermined template. The template may include information that identifies the location where each condition included in the target search conditions (see Figure 8) is set. As a result, the search result acquisition unit 222 may automatically create prompts to be input to the second LLM using the predetermined template and the target search conditions acquired by the search condition acquisition unit 112. The user may also create prompts or modify parts of the prompts created by the search result acquisition unit 222.
[0202] For example, the search result acquisition unit 222 inputs the created prompt to the second LLM. Upon receiving the prompt, the second LLM comprehensively interprets the target search conditions and the second sentence and searches for search targets that match the search conditions from the captured images. Further information such as information accompanying the capture and publicly available information obtained through the internet may be used in interpreting the target search conditions and the second sentence. The second LLM outputs the search results.
[0203] The second LLM may be included in the search result acquisition unit 222. The second LLM may be provided in the information processing device COM4, which is connected to the search result acquisition unit 222 via a communication network NT, as shown in Figure 18. The information processing device COM4 may be the same device as one or more of the information processing devices COM1 to COM3 described above.
[0204] For example, the search result acquisition unit 222 inputs a prompt to the second LVLM and acquires the results of searching for the target from the captured image.
[0205] In detail, if the search result acquisition unit 222 includes a second LLM, the search result acquisition unit 222 inputs a prompt to the second LLM and acquires the result of searching for the target from the captured image using the second sentence.
[0206] If the second LLM is provided in the information processing device COM4, the search result acquisition unit 222 inputs a prompt to the second LLM by sending a prompt to the information processing device COM4. As a result, the second LLM outputs the result of searching for the search target in the captured image using the second sentence. The search result acquisition unit 222 acquires the search result from the information processing device COM4.
[0207] The search result is information indicating that the search target could not be found. If the search target could not be found, the search result acquisition unit 222 may recreate the prompt. The second sentence identification information of the recreated prompt includes the second sentence of the image with the next highest priority after the image used in the previous search. As described above, the selected images are selected in order of priority, starting with those closest to the search starting point. Therefore, the search range can be changed by changing the second sentence of the image used to search for the search target according to the priority of the images.
[0208] Furthermore, as described above, if the search is completed using all of the selected captured images, the image selection unit 121 may re-select captured images in order of priority, starting with those closest to the search starting point, in order to expand the search range.
[0209] The recreated prompt may be the same as the previous prompt, except for the second document identification information. The target search conditions may be relaxed in the recreated prompt. The relaxed target search conditions are as described above.
[0210] The search result acquisition unit 222 inputs the recreated prompt to the second LLM and acquires the search results. The search result acquisition unit 222 repeats the process of recreating the prompt and acquiring the search results using the second LLM until, for example, the same termination conditions as described above are met.
[0211] If the target of the search is found, the search results may include at least one of the following: the location of the target of the search, the direction of movement of the target of the search, the time the image in which the target of the search was found (i.e., the image containing the target of the search) was taken, an image of the target of the search, an image in which the event related to the report was found, and the location and time when the event related to the report occurred. The search results may also include a second sentence, which is a written description of the image in which the target of the search was found. The search results may also include a map image. Note that the information included in the search results is not limited to those exemplified here.
[0212] Figure 19 shows a second example of a prompt and search results.
[0213] The prompt shown in this figure includes "input text" instead of "input image" in the prompt shown in Figure 12. This figure also shows the search results output from the second LLM after the prompt is input to the second LLM. The search results in this figure are examples of information output from the second LLM when the search target is found after inputting the prompt shown in this figure to the second LLM, and are similar to the search results shown in Figure 12.
[0214] (Example of processing operation of the information processing device 200) The information processing performed by the information processing device 200 is represented by a flowchart that is generally similar to the information processing shown in Figure 2. In this embodiment, in step S130, the search unit 220 performs the processing described above and outputs target information related to the search target based on the search results. Except for this point, the information processing according to this embodiment may be the same as the information processing described in Embodiment 1.
[0215] The information processing device 200 may be physically configured in the same way as the information processing device 100 described in Embodiment 1.
[0216] Embodiment 2 has been described above.
[0217] (Function and Effect) According to this embodiment, the search unit 220 obtains the result of searching for a target from the captured image based on a second sentence which is a written description of the contents of the captured image and the search conditions.
[0218] This allows the system to use the second sentence to search for the target within the captured image and output information suitable for responding to the report. Therefore, it becomes possible to respond appropriately to the report.
[0219] According to this embodiment, the search unit 220 inputs the search conditions to the second large-scale language model and obtains the results of searching for the target from the captured image using the second sentence.
[0220] This allows for the use of a large-scale language model to search for targets within captured images and output information suitable for responding to reports. Consequently, it becomes possible to respond appropriately to reports.
[0221] <Embodiment 3> In this embodiment, an example is described in which a summary of the notification audio is output as the first report before outputting the target information.
[0222] As shown in Figure 20, the information processing device 300 according to this embodiment comprises an interpretation unit 310, a search unit 120 similar to that in Embodiment 1, and an output unit 330.
[0223] (Regarding the interpretation unit 310) Similar to the interpretation unit 110 in Embodiment 1, the interpretation unit 310 acquires search conditions for searching for a search target related to the notification voice based on the notification voice. The interpretation unit 310 further acquires a summary text that summarizes the notification voice.
[0224] The interpretation unit 310 inputs the first sentence into, for example, the third large-scale visual language model (third LLM) to obtain a summary. As a result, the interpretation unit 310 obtains a summary that summarizes the voice notification.
[0225] (Detailed example of the interpretation unit 310) As shown in Figure 21, the interpretation unit 310 includes a first text acquisition unit 111 and a search condition acquisition unit 112, similar to those in Embodiment 1. The interpretation unit 310 further includes a summary acquisition unit 313.
[0226] The third LLM is an LLM used for creating summary texts. The third LLM may be an LLM constructed using general techniques. The third LLM may be the same as one or both of the first and second LLMs. The third LLM differs from the first and second LLMs in at least the input prompts.
[0227] The summary acquisition unit 313 inputs the first sentence into the third LLM and acquires a summary of the first sentence. As a result, the summary acquisition unit 313 acquires a summary of the notification audio.
[0228] The third LLM may be included in the summary acquisition unit 313. The summary acquisition unit 313 may be provided in the information processing device COM5, which is connected to each other via a communication network NT, as shown in Figure 22. Note that the information processing device COM5 may be the same device as one or more of the information processing devices COM1 to COM4 described above.
[0229] For example, when the third LLM receives a prompt containing the first sentence, it interprets the first sentence and outputs a summary of the first sentence as a sentence corresponding to the prompt. In this case, the prompt includes, for example, an instruction to create a summary of the first sentence.
[0230] For example, when the 3rd LLM receives a prompt that includes additional information to the report, it interprets the first sentence and the additional information together and outputs a summary of the first sentence as a sentence corresponding to the prompt. In this case, the prompt includes, for example, an instruction to create a summary of the first sentence in accordance with the first sentence and the additional information to the report.
[0231] Publicly available information, such as that obtainable via the internet, may also be used in the interpretation process for creating the summary.
[0232] The prompts entered into the third LLM may be created by the summary acquisition unit 313 or the user according to a predetermined template.
[0233] The template may include information that identifies the location for setting the first sentence. This allows the summary acquisition unit 313 to automatically create a prompt to be entered into the third LLM using a predetermined template and the first sentence acquired by the first sentence acquisition unit 111. The template may further include information that identifies the location for setting the accompanying information for the notification. This allows the summary acquisition unit 313 to automatically create a prompt to be entered into the third LLM using the accompanying information for the notification.
[0234] The user may modify part of the prompt created by the summary acquisition unit 313.
[0235] Figure 23 shows an example of a prompt and summary.
[0236] The prompt in the figure is an example that includes instructions to create a summary of the first sentence, corresponding to the first sentence and accompanying information for the report exemplified in Figure 6. More specifically, the prompt in the figure is an example that includes instructions to create a summary assuming transmission via police radio. The accompanying information for the report includes the time of the report, "hh:mm". For the sake of simplicity, the details of the first sentence in the figure are omitted, but the text data of the first sentence exemplified in Figure 6 is set. The figure also shows an example of a summary output from the third LLM after the prompt is entered into the third LLM.
[0237] (Regarding the output unit 330) The output unit 330 outputs target information based on the search results, similar to the output unit 130 in Embodiment 1.
[0238] The output unit 330 further outputs summary information corresponding to the summary text. The summary information includes, for example, a summary of the first text. The summary information may also include supplementary information for the notification.
[0239] (Example of processing operation of the information processing device 300) The information processing device 300 performs information processing as shown in Figure 24. Information processing is started, for example, when a notification is received. However, the trigger for starting information processing is not limited to receiving a notification.
[0240] Step S111 described above is performed.
[0241] The summary acquisition unit 313 acquires a summary of the first sentence based on the first sentence acquired in step S111 (step S113).
[0242] For example, the summary acquisition unit 313 uses the third LLM to acquire a summary of the first text.
[0243] If the summary acquisition unit 313 includes a third LLM, the summary acquisition unit 313 acquires a summary of the first sentence by inputting a prompt containing the first sentence into the third LLM.
[0244] If the third LLM is provided in the information processing device COM5, the summary acquisition unit 313 inputs the first sentence to the third LLM by sending a prompt containing the first sentence to the information processing device COM5. As a result, the third LLM outputs a summary of the first sentence. The summary acquisition unit 313 acquires the summary from the information processing device COM5.
[0245] The output unit 330 outputs summary information corresponding to the summary text obtained in step S113 (step S330a).
[0246] Steps S112 and S120 described above are performed.
[0247] The output unit 130 outputs target information related to the search target based on the search results, similar to step S130 described above (step S330b).
[0248] The information processing device 300 may be physically configured in the same way as the information processing device 100 described in Embodiment 1.
[0249] Embodiment 2 has been described above.
[0250] (Function and Effects) According to this embodiment, the interpretation unit 310 further acquires a summary text that summarizes the notification voice. The output unit 130 outputs summary information corresponding to the summary text.
[0251] Generally, summarizing text takes less time than searching for a target within captured images. Therefore, a summary of the alert audio can be output as the first report before outputting target information, allowing for early action on the alert audio. Consequently, it becomes possible to respond appropriately to the alert audio.
[0252] According to this embodiment, the interpretation unit 310 inputs the first sentence into the third LLM and obtains a summary.
[0253] Generally, the process of summarizing text by LLM takes less time than the process of VLM or the process of VLM searching for targets within captured images. Therefore, a summary of the alert audio can be output as the first report before outputting target information, allowing for early action on the alert audio. Consequently, it becomes possible to respond appropriately to the alert audio.
[0254] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0255] Furthermore, while the flowcharts used in the above description show multiple steps (processes) in sequence, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impair the content.
[0256] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. An information processing device comprising: an interpretation means for acquiring search conditions for searching for a target related to a notification voice based on the notification voice; a search means for acquiring the result of searching for the target from captured images based on the search conditions; and an output means for outputting target information relating to the target based on the search result. 2. The information processing device according to 1, wherein the target information includes at least one of the location of the target, the direction of movement of the target, the time of capture of the captured image including the target, an image of the target, at least a part of the search conditions, an event related to the notification, the location where the event related to the notification occurred, the time when the event related to the notification occurred, the captured image in which the event related to the notification was discovered, and the notification voice. 3. The information processing device according to 1 or 2, wherein the captured image used for searching for the target is changed if the target cannot be found. 4. The information processing device according to 3, wherein the captured image used for searching for the target is changed based on the search conditions. 5. The search conditions are relaxed if the target cannot be found. An information processing device according to any one of the following. 6. An information processing device according to any one of 1 to 5, wherein the interpretation means includes a first sentence acquisition means for acquiring a first sentence which is a written representation of the content of the notification voice, and a search condition acquisition means for acquiring the search condition based on the first sentence. 7. An information processing device according to any one of 1 to 6, wherein the search means acquires the result of searching for the search target from the captured image based on a second sentence which is a written representation of the content of the captured image and the search condition. 8. An information processing device according to 6 or 7, wherein the interpretation means further acquires a summary sentence which is a summary of the notification voice, and the output means further outputs summary information corresponding to the summary sentence. 9. An information processing device according to 6, wherein the first sentence acquisition means inputs the notification voice to a speech recognition model to acquire the first sentence, and the search condition acquisition means inputs the first sentence to a first large-scale language model to acquire the search condition.10. The information processing device according to any one of 1 to 9, wherein the search means obtains the result of searching for the search target from the captured image using the search conditions and the first large-scale visual language model. 11. The information processing device according to 7, wherein the search means obtains the result of searching for the search target from the captured image using the second sentence, the search conditions and the second large-scale language model. 12. The information processing device according to 8, wherein the interpretation means inputs the first sentence to the third large-scale visual language model to obtain the summary sentence. 13. The information processing device according to 3 or 4, wherein the modification of the captured image used to search for the search target includes at least one of the expansion and movement of the search range. 14. The information processing device according to any one of 1 to 13, wherein the target information is transmitted to one or more personnel terminals held by one or more personnel within a predetermined range from the location of the search target. 15. 1 to 14. An information processing system comprising an information processing device as described in any one of the above, at least one imaging device for acquiring the captured image, and at least one personnel terminal held by a predetermined person. 16. An information processing method comprising at least one computer acquiring search conditions for searching for a search target related to the notification voice based on the notification voice, acquiring the result of searching for the search target from the captured image based on the search conditions, and outputting target information relating to the search target based on the search result. 17. The information processing method according to 16, wherein the target information includes at least one of the location of the search target, the direction of movement of the search target, the time of shooting of the captured image including the search target, the image of the search target, at least a part of the search conditions, the event related to the notification, the location where the event related to the notification occurred, the time when the event related to the notification occurred, the captured image in which the event related to the notification was discovered, and the notification voice. 18. The information processing method according to 16 or 17, wherein the captured image used for searching for the search target is changed if the search target cannot be found. 19. The captured image used for searching for the search target is changed based on the search conditions. The information processing method described above.20. The information processing method according to any one of 16 to 19, wherein the search conditions are relaxed if the search target cannot be found. 21. The information processing method according to any one of 16 to 20, wherein obtaining the search conditions includes obtaining a first sentence which is a written representation of the content of the notification voice, and obtaining the search conditions based on the first sentence. 22. The information processing method according to any one of 16 to 21, wherein obtaining the search results includes obtaining the results of searching for the search target from the captured images based on a second sentence which is a written representation of the content of the captured images and the search conditions. 23. The information processing method according to 21 or 22, wherein obtaining the search conditions further includes obtaining a summary sentence which is a summary of the notification voice, and outputting the target information further includes outputting summary information corresponding to the summary sentence. 24. The information processing method according to 21, wherein the first sentence is obtained by inputting the notification voice into a speech recognition model, and the search conditions are obtained by inputting the first sentence into a first large-scale language model. 25. The information processing method according to any one of 16 to 24, wherein the search results are obtained from the captured images using the search conditions and the first large-scale visual language model. 26. The information processing method according to 22, wherein the search results are obtained from the captured images using the second sentence, the search conditions and the second large-scale language model. 27. The information processing method according to 23, wherein obtaining the search conditions involves inputting the first sentence into the third large-scale visual language model to obtain the summary sentence. 28. The information processing method according to 18 or 19, wherein the modification of the captured images used for searching the search target includes at least one of expanding and moving the search range. 29. The information processing method according to any one of 16 to 28, wherein the target information is transmitted to one or more personnel terminals held by one or more personnel within a predetermined range from the location of the search target. 30. At least one computer, 16 to 29. A program for executing one of the information processing methods described in any one of the following.31. A recording medium on which a program is stored that causes at least one computer to perform any one of the information processing methods described in 16. to 29.
[0257] This application claims priority based on Japanese Patent Application No. 2025-010363, filed on 24 January 2025, and incorporates all of its disclosures herein.
[0258] 100, 200, 300 Information processing device 110, 310 Interpretation unit 111 First document acquisition unit 112 Search condition acquisition unit 120, 220 Search unit 130, 330 Output unit 240 Image document conversion device 313 Summary acquisition unit DB1 Image storage device DB2 Second document storage device
Claims
1. An information processing device comprising: an interpretation means for obtaining search conditions for searching for a target related to a report voice based on the report voice; a search means for obtaining the results of searching for the target from captured images based on the search conditions; and an output means for outputting target information related to the target based on the search results.
2. The information processing device according to claim 1, wherein the target information includes at least one of the following: the location of the search target, the direction of movement of the search target, the time of taking the captured image including the search target, the image of the search target, at least a part of the search conditions, the event related to the notification, the location where the event related to the notification occurred, the time when the event related to the notification occurred, the captured image in which the event related to the notification was discovered, and the notification audio.
3. The information processing apparatus according to claim 1 or 2, wherein the captured image used for searching for the target to be searched is changed if the target to be searched cannot be found.
4. The information processing device according to claim 3, wherein the captured image used for searching for the target to be searched is changed based on the search conditions.
5. The information processing apparatus according to any one of claims 1 to 4, wherein the search conditions are relaxed if the search target cannot be found.
6. The information processing apparatus according to any one of claims 1 to 5, wherein the interpretation means includes a first sentence acquisition means for acquiring a first sentence which is a written representation of the content of the notification voice, and a search condition acquisition means for acquiring the search conditions based on the first sentence.
7. The information processing apparatus according to any one of claims 1 to 6, wherein the search means obtains the result of searching for the search target from the captured image based on a second sentence which is a written representation of the contents of the captured image and the search conditions.
8. The information processing apparatus according to claim 6 or 7, wherein the interpretation means further obtains a summary text that summarizes the notification voice, and the output means further outputs summary information corresponding to the summary text.
9. The information processing apparatus according to claim 6, wherein the first text acquisition means inputs the notification voice to a speech recognition model to acquire the first text, and the search condition acquisition means inputs the first text to a first large-scale language model to acquire the search condition.
10. The information processing apparatus according to any one of claims 1 to 9, wherein the search means obtains the result of searching for the search target from the captured image using the search conditions and the first large-scale visual language model.
11. The information processing apparatus according to claim 7, wherein the search means obtains the result of searching for the search target from the captured image using the second sentence, the search conditions, and the second large-scale language model.
12. The information processing apparatus according to claim 8, wherein the interpretation means inputs the first text into a third large-scale visual language model to obtain the summary text.
13. The information processing apparatus according to claim 3 or 4, wherein the modification of the captured image used for searching for the search target includes at least one of expanding and moving the search range.
14. The information processing device according to any one of claims 1 to 13, wherein the target information is transmitted to one or more personnel terminals held by one or more personnel within a predetermined range from the location of the search target.
15. An information processing method comprising: at least one computer obtaining search conditions for searching for a target related to a notification voice based on the notification voice; obtaining the result of searching for the target from captured images based on the search conditions; and outputting target information related to the target based on the search result.
16. A recording medium on which a program is recorded that causes at least one computer to perform the following actions: to acquire search conditions for searching for a target related to a notification voice based on the notification voice; to acquire the results of searching for the target from captured images based on the search conditions; and to output target information related to the target based on the search results.