Information processing systems and programs

The information processing system addresses the challenge of balancing privacy and emergency response by adjusting metadata disclosure based on location and urgency, enabling effective communication during emergencies.

JP2026081433APending Publication Date: 2026-05-19FUJIFILM BUSINESS INNOVATION CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FUJIFILM BUSINESS INNOVATION CORP
Filing Date
2024-11-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Network cameras face challenges in balancing the protection of personal information and the need for rapid response during emergencies, as removing private information from images can lead to insufficient data during disasters, while transmitting images without removal risks privacy leakage.

Method used

An information processing system that determines the scope of disclosure based on the camera's installation location and the urgency of the situation, adjusting the inclusion of personally identifiable information in metadata, and uses a large-scale language model to convert image information into text or audio when transmission fails.

Benefits of technology

The system ensures both personal information protection and rapid response by dynamically controlling the disclosure of metadata, allowing for wider information sharing during emergencies and local notification when transmission fails.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026081433000001_ABST
    Figure 2026081433000001_ABST
Patent Text Reader

Abstract

When transferring information extracted from images continuously captured by cameras installed in designated locations to an external source, the system ensures both the protection of personal information and ease of response in the event of an emergency, even if the captured images include images of people. [Solution] The information processing system of this disclosure includes a processor, which determines the scope of disclosure to determine the extent to which personally identifiable information should be included in the information extracted from the captured images, based on the level of protection of personal information according to the installation location of a camera that continuously captures images of a predetermined location and the urgency of the abnormal situation that has occurred, and transmits the information extracted from the images captured by the camera, including personally identifiable information, to an external device according to the determined scope of disclosure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , , , ,

[0006] , , , ,

[0005] , , , , , , , , ,

[0001] The present disclosure relates to an information processing system and a program.

Background Art

[0002] Patent Document 1 discloses a video surveillance system that can monitor a person's actions while protecting their privacy by identifying the person's portion in an image of the person to be monitored, synthesizing the abstracted image of the identified person's portion and the original image, and then transmitting the synthesized image.

[0003] Patent Document 2 discloses a position information system that can display information of all mobile terminals in an emergency while protecting privacy under normal circumstances.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] In recent years, network cameras have been used for detecting congestion in restaurants, monitoring in unmanned checkout stores, detecting intruders for crime prevention, detecting people's falls or accidents, etc. Such network cameras are configured to be installed at preset locations and transmit continuously captured images to an external server. Furthermore, in recent years, an endpoint camera has also been proposed that performs image analysis by AI (Artificial Intelligence) processing on the camera side and transmits image data with privacy information excluded to an external device when an image of a person is included in the captured image.

[0006] However, if such endpoint cameras always remove private information from captured images before transmitting them externally, there is a problem in that in the event of an emergency such as a disaster or accident, there will be insufficient information to respond quickly. On the other hand, if such endpoint cameras always transmit images without removing private information externally, there is a concern about the leakage of private information, and individual privacy cannot be adequately protected.

[0007] The purpose of this disclosure is to provide an information processing system and program that can simultaneously protect personal information and facilitate responses in the event of an abnormal situation when transferring information extracted from images continuously captured by cameras installed in pre-set locations to an external source, even if the captured images include images of people. [Means for solving the problem]

[0008] An information processing system in the first aspect of this disclosure comprises a processor which determines the scope of disclosure to which personally identifiable information should be included in the information extracted from the captured images, based on the level of protection of personal information corresponding to the installation location of a camera that continuously captures images of a predetermined location and the urgency of the abnormal situation that has occurred. Depending on the determined scope of disclosure, the information extracted from the images captured by the camera, including personally identifiable information, is transmitted to an external device.

[0009] The information processing system in the second aspect of this disclosure further comprises a memory in the information processing system of the first aspect, the memory storing the scope of disclosure in association with a combination of the level of protection of personal information according to the location where the camera is installed and the urgency of the abnormal situation that has occurred. The processor determines the scope of disclosure using the level of personal information protection appropriate to the camera's installation location and the urgency of the abnormal situation that occurred.

[0010] The information processing system in the third aspect of this disclosure is an information processing system in the second aspect, wherein the processor is configured such that the scope of disclosure is narrower when the level of protection of the personal information is higher, and wider when the urgency of the abnormal situation is higher.

[0011] The information processing system of the fourth aspect of this disclosure is set such that, in the information processing system of the second aspect, the scope of disclosure expands in stages in the following order as the range of personal information to be included in the information extracted from the image captured by the camera expands: skeletal information of a person obtained by estimating the skeleton of a person in the captured image, identification information of only persons who have given prior consent to the provision of personal information, and identification information of both persons who have given prior consent and persons who have not given consent to the provision of personal information.

[0012] The fifth aspect of the information processing system of the present disclosure, in the information processing system of the first aspect, when an abnormal situation occurs and the information extracted from the captured image cannot be transferred to a pre-set transfer destination, the processor inputs the information extracted from the image captured by the camera to a large-scale language model that converts the input information into text information and outputs it, thereby obtaining text information regarding the image content of the captured image. The acquired text information is output as audio information via the audio output unit.

[0013] The information processing system in the sixth aspect of this disclosure, in the information processing system in the fifth aspect, the processor performs skeletal estimation of a person in a captured image to detect whether or not there is a postural abnormality in the person, When an abnormal situation occurs, and information extracted from the captured image cannot be transferred to a pre-configured destination, and if a posture abnormality of a person is detected in the image, the information extracted from the image in which the posture abnormality of the person was detected is input to the large-scale language model to obtain text information regarding the state of the person in the captured image. The acquired text information is output as audio information via the audio output unit.

[0014] The seventh aspect of this disclosure includes a step of determining the scope of disclosure, which determines the extent to which personally identifiable information should be included in the information extracted from the captured images, based on the level of personal information protection corresponding to the location of a camera that continuously captures images of a predetermined location and the urgency of the abnormal situation that occurs, Depending on the determined scope of disclosure, the computer is instructed to perform the following steps: transmit information extracted from images captured by the camera, including personally identifiable information, to an external device. [Effects of the Invention]

[0015] According to the information processing system of the first aspect of this disclosure, when transferring information extracted from images continuously captured by a camera installed in a predetermined location to an external location, it is possible to achieve both the protection of personal information and ease of response in the event of an abnormal situation, even if the captured images include images of people.

[0016] According to the information processing system of the second aspect of this disclosure, the scope of disclosure can be set by combining the level of protection of personal information and the urgency of the abnormal situation that occurred.

[0017] According to the information processing system of the third aspect of this disclosure, it is possible to set the scope of disclosure in stages by combining the level of protection of personal information and the urgency of the abnormal situation.

[0018] According to the information processing system of the fourth aspect of the present disclosure, it is possible to increase the amount of information transmitted to an external device as the range of personal information included in the information extracted from the captured image becomes wider.

[0019] According to the information processing system of the fifth aspect of the present disclosure, even when the information extracted from the image captured when an abnormal situation occurs cannot be transferred to a preset transfer destination, it is possible to notify the content of the image captured by a person existing around the camera.

[0020] According to the information processing system of the sixth aspect of the present disclosure, even when the information extracted from the image captured when an abnormal situation occurs cannot be transferred to a preset transfer destination, it is possible to notify a person existing around the camera that an abnormal posture has occurred to a person in the image.

[0021] According to the program of the seventh aspect of the present disclosure, when transferring the information extracted from the images continuously captured by a camera installed at a preset location to the outside, even when the captured image includes a person's image, it is possible to achieve both the protection of personal information and the ease of dealing with abnormal situations.

Brief Description of Drawings

[0022] [Figure 1] It is a diagram showing the system configuration of the information processing system of an embodiment of the present disclosure. [Figure 2] It is a block diagram showing the hardware configuration of the camera 10 in an embodiment of the present disclosure. [Figure 3] It is a block diagram showing the functional configuration of the camera 10 in an embodiment of the present disclosure. [Figure 4] It is a diagram showing an example of the urgency table stored in the table information storage unit 34. <( [Figure 5] It is a diagram showing an example of the privacy consideration level table stored in the table information storage unit 34. [Figure 6]This figure shows an example of a metadata generation policy table stored in the table information storage unit 34. [Figure 7] This figure shows an example of a metadata content determination table stored in the table information storage unit 34. [Figure 8] This diagram illustrates a specific example of metadata generated by the metadata generation unit 36. [Figure 9] This flowchart shows the overall operation of an information processing system according to one embodiment of the present disclosure. [Figure 10] This diagram shows how metadata is generated based on the urgency level and the location of camera 10. [Figure 11] This figure shows the hardware configuration of camera 10A, which is a modified example of one embodiment of the present disclosure. [Figure 12] This is a block diagram showing the functional configuration of camera 10A, which is a modified example of one embodiment of the present disclosure. [Figure 13] This flowchart shows the overall operation of the processing system in a modified embodiment of one embodiment of the present disclosure. [Modes for carrying out the invention]

[0023] Next, embodiments of the present disclosure will be described in detail with reference to the drawings.

[0024] Figure 1 shows the system configuration of an information processing system according to one embodiment of the present disclosure.

[0025] An information processing system according to one embodiment of the present disclosure, as shown in Figure 1, consists of a camera 10 installed in a space to be monitored and a management server 20 connected to the camera 10 by a network such as the Internet 30.

[0026] Camera 10 is installed in spaces to be monitored, such as stores, offices, factories, hospitals, and elderly care facilities. Camera 10 is used for purposes such as detecting congestion levels within the installed space, monitoring unmanned payment stores, detecting intruders for security purposes, and detecting falls or accidents.

[0027] Here, camera 10 is a so-called endpoint camera, and is configured to perform AI (Artificial Intelligence) processing on the captured images to analyze them and send metadata, from which privacy information has been removed, to the management server 20. Camera 10 is installed in a pre-configured location, continuously captures images of the monitored space, and transfers the information extracted from the captured images to the management server 20, which is an external device.

[0028] The management server 20 monitors the space where the camera 10 is installed by receiving metadata from the camera 10 via the internet 30, from which privacy information has been removed.

[0029] However, if camera 10 always sends information to the management server 20 with privacy information removed from the captured images, there is a problem in that in the event of an emergency such as a disaster or accident, the management server 20 will not have enough information to respond quickly.

[0030] Therefore, in this embodiment, the information processing system is designed to ensure both the protection of personal information and ease of response in the event of an abnormal situation, even when the captured image includes images of people, by performing the control described below.

[0031] Next, Figure 2 shows the hardware configuration of the camera 10 in the information processing system of this embodiment.

[0032] As shown in Figure 2, the camera 10 includes a CPU 11, memory 12, a storage device 13 such as flash memory, a communication interface (IF) 14 for transmitting and receiving data to and from external devices via a network such as the Internet 30, an AI processor 15, and a shooting unit 16. These components are connected to each other via a control bus.

[0033] The CPU 11 is a processor that controls the operation of the camera 10 by executing predetermined processes based on a control program stored in the memory 12 or storage device 13. In this embodiment, the CPU 11 is described as reading and executing a control program stored in the memory 12 or storage device 13, but it is not limited to this. This control program may be provided in the form of a computer-readable recording medium. For example, this program may be provided in the form of a CD (Compact Disc)-ROM and DVD (Digital Versatile Disc)-ROM recorded on an optical disc, or in the form of a USB (Universal Serial Bus) memory and memory card recorded on a semiconductor memory. Furthermore, this control program may be acquired from an external device via a communication line connected to the communication interface 14. In addition, this control program may be provided as a standalone application software, or it may be incorporated into the software of each device as a function of the camera 10.

[0034] The imaging unit 16 continuously captures images of the monitored space using an image sensor such as a CCD (Charge Coupled Device) or CMOS (Complementary Metal-Oxide-Semiconductor) image sensor. Continuously capturing images here includes both capturing moving images and capturing still images at regular intervals.

[0035] The AI ​​processor 15 is a processor that performs AI image analysis using a neural network on images captured by the shooting unit 16. The AI ​​processor 15 uses various AI models to recognize objects contained in the image, identify people through face recognition, and detect people's posture through skeletal estimation. Specifically, the AI ​​processor 15 may, for example, use an object detection algorithm such as YOLO as an AI model to recognize objects, or use a skeletal estimation algorithm such as HRNet to detect people's posture.

[0036] Furthermore, a disaster detection sensor 17 located outside the camera 10 detects the occurrence of disasters such as earthquakes and fires, and the detection results are notified to the CPU 11. The disaster detection sensor 17 is, for example, an earthquake sensor and a fire sensor, and detects when a fire or earthquake occurs in the monitored space.

[0037] Figure 3 is a block diagram showing the functional configuration of the camera 10 realized by the execution of the control program described above.

[0038] As shown in Figure 3, the camera 10 of this embodiment includes an imaging unit 31, an image analysis unit 32, an urgency determination unit 33, a table information storage unit 34, a metadata content determination unit 35, a metadata generation unit 36, and a metadata transmission unit 37.

[0039] The imaging unit 31 is composed of the imaging unit 16 described above and continuously captures images of a pre-set location to be monitored.

[0040] The image analysis unit 32 is composed of the AI ​​processor 15 described above and performs image analysis processing such as object detection and human skeleton estimation on the image captured by the imaging unit 31.

[0041] The urgency determination unit 33 determines the urgency of an abnormal situation, such as an earthquake or fire, when it occurs. Specifically, the urgency determination unit 33 determines the urgency of the abnormal situation based on the analysis results from the image analysis unit 32 and the disaster detection results from the disaster detection sensor 17. This abnormal situation may include not only the occurrence of disasters such as earthquakes and fires, but also things like people falling in the monitored space. The urgency is determined in four stages, for example, from level 0 to level 3. Details of this urgency will be described later. In addition, even if the occurrence of a disaster such as a fire or earthquake is detected, the urgency determination unit 33 will determine that the urgency is higher if a person who has fallen and is not moving is also detected.

[0042] The table information storage unit 34 stores various types of table information as shown in Figures 4 to 7.

[0043] The urgency level table shown in Figure 4 illustrates the correspondence between the nature of the abnormal condition that occurred and its urgency level. Here, the urgency level is set in four stages, from level 0 to level 3. In a normal state where no disaster has occurred, the urgency level is set to level 0. If a disaster occurs that does not require immediate evacuation, the urgency level is set to level 1. If a disaster occurs that requires immediate evacuation, the urgency level is set to level 2. Furthermore, if a disaster occurs that requires immediate evacuation, and an immobile person is detected, the urgency level is set to level 3.

[0044] The privacy consideration level table shown in Figure 5 illustrates the correspondence between privacy considerations and privacy consideration levels. The privacy consideration level is a level of protection for personal information that sets the degree to which an individual's privacy is considered, and is set in four stages: "High," "Medium," "Low," and "None." When the privacy consideration level is "High," maximum consideration is given to privacy, and information that can identify an individual will not be distributed to anyone who has not given prior consent, under any circumstances. When the privacy consideration level is "Medium," privacy is considered, and metadata containing personally identifiable information will not be distributed as much as possible. When the privacy consideration level is "Low," personally identifiable information will be distributed to anyone who has not given prior consent in emergencies. Furthermore, when the privacy consideration level is "None," privacy is not considered much, and detailed image data will be distributed in addition to metadata.

[0045] Furthermore, individuals entering or leaving the space being filmed are required to indicate in advance whether or not they consent to their personal information being included in metadata and distributed externally. Therefore, in cases where the urgency level is 0 and no disaster or other emergency has occurred, personal information of individuals who have not given prior consent will not be included in the metadata distributed externally.

[0046] The metadata generation policy table shown in Figure 6 is a table that shows the correspondence between camera installation locations and privacy consideration levels. Referring to the table in Figure 6, for example, if the camera is installed in the north area of ​​the office, the privacy consideration level is set to "medium". Similarly, if the camera is installed in the south area of ​​the office, a public space, or around the restrooms, the privacy consideration levels are set to "low", "none", and "high", respectively. In other words, a privacy consideration level that takes into account the characteristics of the location where camera 10 is installed is pre-set for each installation location.

[0047] Furthermore, the table information storage unit 34 stores a metadata content determination table as shown in Figure 7. This metadata content determination table associates metadata content with each combination of the privacy consideration level according to the camera installation location and the urgency of the abnormal situation that occurred.

[0048] Here, the metadata content specifically refers to the scope of disclosure, which determines to what extent personally identifiable information is included in the metadata and other information extracted from the images captured by the imaging unit 31.

[0049] For example, referring to the decision table in Figure 7, when the urgency level is 0 and the privacy consideration level is "low," the metadata is set to include only skeletal information. However, even when the privacy consideration level is "low," if the urgency level is 1, the metadata is set to include not only skeletal information but also the ID information of individuals who have given prior consent. Furthermore, even when the privacy consideration level is "low," if the urgency level is 2, the metadata is set to include not only skeletal information but also the ID information of all individuals, both those who have given prior consent and those who have not. Moreover, even when the privacy consideration level is "low," if the urgency level is 3, the metadata is set to include skeletal information and the ID information of all individuals, as well as partial images of the images captured by the imaging unit 31.

[0050] Here, a person's ID information refers to information that can identify that person, such as their name, employee number, identification number, and other types of information.

[0051] The decision table shown in Figure 7 reveals that even for cameras installed in the same location, i.e., cameras with the same level of privacy consideration, the scope of personal information disclosed in the metadata expands as the urgency increases. In other words, even if the settings are configured to minimize the disclosure of personal information in normal, low-urgency situations, the settings increase as the urgency increases, allowing for the disclosure of personal information in the metadata.

[0052] The metadata content determination unit 35 determines the scope of disclosure, which is the extent to which personally identifiable information should be included in the metadata extracted from the captured images, based on the level of privacy considerations corresponding to the installation location of the camera 10 and the urgency of the abnormal situation that occurred.

[0053] Here, the metadata content determination unit 35 determines the scope of disclosure of personal information in the metadata based on the determination table shown in Figure 7, using the level of privacy consideration corresponding to the installation location of the camera 10 and the urgency of the abnormal situation that occurred.

[0054] As shown in Figure 7, the scope of disclosure is set such that the higher the level of privacy consideration, the narrower the range of personal information included in the metadata extracted from the images captured by camera 10, and the higher the urgency of the abnormal situation, the wider the range of personal information included in the metadata extracted from the images captured by camera 10.

[0055] Furthermore, the scope of disclosure is set to expand in stages, in the following order, as the range of personal information to be included in the metadata extracted from the images captured by camera 10 increases: the skeletal information of a person obtained by estimating the person's skeleton in the captured image, the ID information of only those who have given prior consent to the provision of personal information, and the ID information of both those who have given prior consent and those who have not.

[0056] The metadata generation unit 36 ​​generates metadata that includes personally identifiable information in the metadata extracted from the image captured by the camera 10's shooting unit 31, according to the disclosure scope determined by the metadata content determination unit 35.

[0057] The metadata transmission unit 37 transmits the metadata generated by the metadata generation unit 36 ​​to the management server 20, which is an external device.

[0058] Next, a specific example of metadata generated by the metadata generation unit 36 ​​will be explained with reference to Figure 8.

[0059] First, the image analysis unit 32 performs image analysis on the image captured by the shooting unit 31, including object detection, estimation of human skeletons, and identification of individuals through facial recognition. Then, the metadata generation unit 36 ​​generates metadata, such as skeletal information and recognition angle information for each person in the image, based on the analysis results from the image analysis unit 32. Figure 8 shows a case where only skeletal information is generated as metadata, without including personally identifiable information. Specifically, this skeletal information consists of coordinate information for the face, eyes, shoulders, waist, hands, and feet of the target person.

[0060] In this embodiment, we describe a case where the information to be included in the metadata is information about a person in the image, but information about an object in the image may also be included as metadata.

[0061] Thus, metadata may include the following types of information. (1) Object information in the image Coordinates, names, and recognition accuracy information within space (2) Skeletal information of the person in the image Coordinate points of the face, eyes, shoulders, waist, hands, and feet, and recognition accuracy information. (3) Personal identification information Name of a person Images containing people (images where everything except the subject is obscured)

[0062] Next, the operation of the information processing system of this embodiment will be described in detail with reference to the drawings.

[0063] Figure 9 is a flowchart showing the overall operation of the information processing system in this embodiment.

[0064] First, in step S101, the imaging unit 31 captures an image of the space to be monitored. Then, in step S102, the image analysis unit 32 performs image analysis processing such as object detection, skeletal estimation, and face recognition on the captured image.

[0065] Next, in step S103, the urgency determination unit 33 determines the current urgency and sets the urgency level based on the urgency table shown in Figure 4.

[0066] Then, in step S104, the metadata content determination unit 35 determines the content of the metadata to be generated according to the level of urgency set in the urgency determination unit 33 and the privacy consideration level based on the location where the camera 10 is installed. For example, if the location where the camera 10 is installed is a public space, the metadata content determination unit 35 refers to the metadata generation policy table shown in Figure 6 and determines that the privacy consideration level is "none". Then, based on the determined privacy consideration level and the level of urgency set in the urgency determination unit, the metadata content determination unit 35 refers to the metadata content determination table shown in Figure 7 and determines the content of the metadata to be generated.

[0067] Then, in step S105, the metadata generation unit 36 ​​generates metadata based on the metadata content determined by the metadata content determination unit 35, by referring to the analysis results from the image analysis unit 32.

[0068] Finally, in step S106, the metadata transmission unit 37 transmits the metadata generated by the metadata generation unit 36 ​​to an external device such as the management server 20.

[0069] Figure 10 shows how metadata is generated based on the urgency level and the location of camera 10.

[0070] Figure 10 illustrates an example where the urgency level is set to level 2 and camera 10 is installed in the north area of ​​the office. It also explains that the captured image shows three individuals, users A, B, and C, and that only user A has consented to the disclosure of their personal information.

[0071] In this case, the metadata content determination unit 35 determines that the privacy consideration level is "medium" based on the metadata generation policy table in Figure 6, since the camera 10 is installed in the north area of ​​the office. Then, referring to the content determination table in Figure 7, the metadata content determination unit 35 determines that the metadata to be generated when the privacy consideration level is "medium" and the urgency level is 2 is skeletal information and consenter ID information.

[0072] Therefore, based on the analysis results by the image analysis unit 32, the metadata generation unit 36 ​​generates metadata that includes not only the skeletal information of users A to C in the image, but also the name information, which is the ID information of user A who has given prior consent to the disclosure of personal information. Referring to Figure 10, it can be seen that the generated metadata includes the name of user A, "Taro Yamada".

[0073] [Differentiation] Next, we will describe camera 10A, which is a modified version of camera 10 in this embodiment.

[0074] Since the information processing system of this embodiment is intended for use during disasters, it is anticipated that network and other equipment failures may occur as a result of the disaster. Furthermore, if such network or other equipment failures occur, it may become impossible to transfer the generated metadata to the management server 20. Therefore, in a modified version of this embodiment, if a situation arises where the metadata cannot be transferred to the management server 20, a function is implemented to notify people around the camera of the metadata content via an audio output means such as a speaker.

[0075] Figure 11 shows the hardware configuration of camera 10A, which is a modified example of this embodiment. The hardware configuration of camera 10A shown in Figure 11 is the same as that of camera 10 shown in Figure 2, but with the addition of a speech synthesis unit 18, a speaker 19, and a large language model (hereinafter abbreviated as LLM (Large Language Models)) 40. In Figure 11, the same reference numerals are used for components that are the same as those in Figure 2, and their descriptions are omitted.

[0076] The speech synthesis unit 18 converts the generated text information into an audio signal by performing speech synthesis processing. The speaker 19 outputs the audio signal generated by the speech synthesis unit 18 to the outside as audio.

[0077] The LLM40 has the function of converting various information, such as input images and metadata, into text information and outputting it. For example, by inputting an image of a person who has fallen, and metadata extracted from this image, into the LLM40, this information is converted into text information such as "A person has fallen."

[0078] Next, Figure 12 shows the functional configuration of camera 10A, which is a modified example of this embodiment. The functional configuration of camera 10A shown in Figure 12 is the same as that of camera 10 shown in Figure 3, but with the addition of an audio conversion unit 38 and an audio output unit 39. In Figure 12, the same reference numerals are used for components that are the same as those in Figure 3, and their descriptions are omitted.

[0079] When an abnormal situation occurs and the metadata cannot be transferred to the management server 20, which is a pre-configured transfer destination, the voice conversion unit 38 inputs the metadata received from the metadata generation unit 36 ​​into the LLM 16 shown in Figure 11, thereby obtaining text information about the image content of the captured image. The text information obtained by the voice conversion unit 38 is then output as audio information via the audio output unit 39.

[0080] For example, if the image analysis unit 32 estimates the skeleton of a person in an image captured by the imaging unit 31 and detects whether or not there is a postural abnormality in that person, and a postural abnormality such as a person falling is detected in the image, and the metadata extracted from the captured image cannot be transferred to a pre-set transfer destination, the voice conversion unit 38 inputs the information extracted from the image in which the postural abnormality of the person was detected into the LLM 40. The voice conversion unit 38 then obtains text information about the state of the person in the captured image from the LLM 40 and outputs the obtained text information as voice information via the voice output unit 39.

[0081] The operation in a modified version of this embodiment will be explained with reference to the flowchart in Figure 13. Note that the flowchart in Figure 13 differs from the flowchart in Figure 9 only in that steps S201 and S202 have been added. Therefore, in the following explanation, only steps S201 and S202 will be described.

[0082] In the modified camera 10A of this embodiment, in step S201, it is determined whether the communication equipment is functioning correctly by whether or not metadata can be transferred to the management server 20. If it is determined in step S201 that the communication equipment is functioning correctly, the generated metadata is sent to the management server 20 by the metadata transmission unit 37.

[0083] Then, if it is determined in step S201 that the communication equipment is not functioning correctly, in step S202, the generated metadata is output as audio by the voice conversion unit 38 and the voice output unit 39. For example, if the metadata includes the name information of the consentee, an audio message such as "Mr. A has fallen" is output from the camera 10A to the surroundings.

[0084] In this embodiment, each process is executed on any computer. Furthermore, any computer may execute these processes using a processor as hardware, a program as software, or a combination thereof. In that case, the processor is configured to work in cooperation with the program to execute the various processes in this embodiment, and can function as a unit or means in this embodiment. Also, the execution order of the processes by the processor is not limited to the order described and may be changed as appropriate. Any computer may be a general-purpose computer, a computer designed for a specific purpose, a workstation, or any other system capable of executing each process.

[0085] A processor may consist of one or more hardware components, and the type of hardware is not limited. For example, a processor may consist of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a programmable logic device such as an FPGA (Field Programmable Gate Array), a dedicated circuit for executing a specific process such as an ASIC (Application Specific Integrated Circuit), a GPU (Graphic Processing Unit), or an NPU (Neural Processing Unit). Furthermore, the type of hardware may be a combination of different types of hardware. When multiple hardware components are configured to execute one or more processes of a processor, these components may reside in physically separate devices or in the same device. Also, in any embodiment, the order of each process performed by the processor is not limited to the order described above and may be changed as appropriate. Hardware is composed of electrical circuits (circuitry) that combine circuit elements such as semiconductor elements.

[0086] Furthermore, the program may be firmware or software such as microcode. Alternatively, the program may be, for example, a group of program modules, each function of which may be implemented by a processor configured to perform its respective function. The program may be program code or multiple code segments stored on one or more non-temporary computer-readable media (e.g., storage media or other storage). The program may be divided and stored on multiple non-temporary computer-readable media located on physically separate devices. The program code or code segments may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. The program code or code segments may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents. The program of this application may also be provided as a program product.

[0087] Furthermore, the processor operations in each of the above embodiments may not be performed by a single processor, but may also be performed by multiple processors located in physically separate locations working together. Also, the order of the processor operations is not limited to the order described in each of the above embodiments, and may be changed as appropriate.

[0088] In this embodiment, "system" includes both systems composed of multiple devices and systems composed of a single device.

[0089] [Differentiation] The above embodiment described an application of the present disclosure to an endpoint camera that performs AI processing on images captured inside the camera 10. However, the present disclosure is not limited to such a configuration and can be similarly applied to configurations in which AI processing is performed and image analysis is carried out outside the camera, which is the imaging device.

[0090] [Note] (((1))) Equipped with a processor, The aforementioned processor, Based on the level of personal information protection required for cameras that continuously capture images of pre-set locations, and the urgency of any abnormal situations that occur, the scope of disclosure is determined, which determines to what extent personally identifiable information should be included in the information extracted from the captured images. Depending on the determined scope of disclosure, information extracted from images captured by the camera, including personally identifiable information, is transmitted to an external device. Information processing system.

[0091] (((2))) With even more memory, The memory stores the scope of disclosure in association with a combination of the level of protection of personal information corresponding to the installation location of the camera and the urgency of the abnormal situation that occurred. The processor determines the scope of disclosure using the level of personal information protection corresponding to the camera's installation location and the urgency of the abnormal situation that occurred. The information processing system described in (((1))).

[0092] (((3))) The scope of disclosure is set such that the higher the level of protection of the personal information, the narrower the range of personal information included in the information extracted from the images captured by the camera, and the higher the urgency of the abnormal situation, the wider the range of personal information included in the information extracted from the images captured by the camera. The information processing system described in (((2))).

[0093] (((4))) The scope of disclosure is set to expand in stages, in the following order, as the range of personal information to be included in the information extracted from the image captured by the camera expands: skeletal information of a person obtained by estimating the skeleton of a person in the captured image; identification information of only persons who have given prior consent to the provision of personal information; and identification information of both persons who have given prior consent and persons who have not. The information processing system described in (((2))).

[0094] (((5))) When an abnormal situation occurs and the information extracted from the captured image cannot be transferred to a pre-configured destination, the processor inputs the information extracted from the image captured by the camera into a large-scale language model that converts the input information into text information and outputs it, thereby obtaining text information about the image content of the captured image. The acquired text information is output as audio information via the audio output unit. An information processing system described in any one of (((1))) through (((4))).

[0095] (((6))) The processor performs skeletal estimation of a person in the captured image to detect whether or not there is a postural abnormality in the person. When an abnormal situation occurs, and information extracted from the captured image cannot be transferred to a pre-configured destination, and if a posture abnormality of a person is detected in the image, the information extracted from the image in which the posture abnormality of the person was detected is input to the large-scale language model to obtain text information regarding the state of the person in the captured image. The acquired text information is output as audio information via the audio output unit. The information processing system described in (((5))).

[0096] (((7))) The process involves determining the scope of disclosure, which is determined based on the level of personal information protection required for cameras that continuously capture images of pre-set locations, and the urgency of any abnormal situations that may occur, and deciding to what extent personally identifiable information should be included in the information extracted from the captured images. Depending on the determined scope of disclosure, the process includes the step of transmitting information extracted from images captured by the camera, including personally identifiable information, to an external device, A program that causes a computer to execute something.

[0097] According to the information processing system (((1))), when transferring information extracted from images continuously captured by cameras installed in pre-set locations to an external source, it is possible to achieve both the protection of personal information and ease of response in the event of an abnormal situation, even if the captured images include images of people.

[0098] According to the information processing system (((2))), the scope of disclosure can be set by combining the level of protection of personal information and the urgency of the abnormal situation that occurred.

[0099] According to the information processing system in (((3))), the scope of disclosure can be set in stages by combining the level of protection of personal information and the urgency of the abnormal situation.

[0100] According to the information processing system (((4))), the wider the range of personal information to be included in the information extracted from the captured image, the greater the amount of information that can be transmitted to an external device.

[0101] According to the information processing system (((5))), even if it is not possible to transfer the information extracted from the captured image to a pre-set destination when an abnormal situation occurs, it is possible to notify people in the vicinity of the camera of the contents of the captured image.

[0102] According to the information processing system (((6))), even if it is not possible to transfer the information extracted from the image taken when an abnormal situation occurs to a pre-set destination, it is possible to notify people in the vicinity of the camera that a posture abnormality has occurred in the person in the image.

[0103] According to the program in (((7))), when transferring information extracted from images continuously captured by cameras installed in pre-set locations to an external source, it is possible to achieve both the protection of personal information and ease of response in the event of an abnormal situation, even if the captured images include images of people. [Explanation of symbols]

[0104] 10, 10A Camera 11 CPU 12 memory 13 Storage device 14. Communication Interface 15 AI Processors 16 Filming Unit 17. Disaster detection sensors 18. Speech Synthesis Unit 19 speakers 20 Management Server 30 Internet 31 Photography Department 32 Image Analysis Department 33 Urgency Judgment Department 34 Table Information Storage Unit 35 Metadata Content Determination Unit 36 Metadata Generation Unit 37 Metadata transmission unit 38. Voice Conversion Unit 39 Audio output section 40. Large-Scale Language Models (LLMs)

Claims

1. Equipped with a processor, The aforementioned processor, Based on the level of personal information protection required for cameras that continuously capture images of pre-set locations, and the urgency of any abnormal situations that occur, the scope of disclosure is determined, which determines to what extent personally identifiable information should be included in the information extracted from the captured images. Depending on the determined scope of disclosure, information extracted from images captured by the camera, including personally identifiable information, is transmitted to an external device. Information processing system.

2. With even more memory, The memory stores the scope of disclosure in association with a combination of the level of protection of personal information corresponding to the installation location of the camera and the urgency of the abnormal situation that occurred. The processor determines the scope of disclosure using the level of personal information protection corresponding to the camera's installation location and the urgency of the abnormal situation that occurred. The information processing system according to claim 1.

3. The scope of disclosure is set such that the higher the level of protection of the personal information, the narrower the range of personal information included in the information extracted from the images captured by the camera, and the higher the urgency of the abnormal situation, the wider the range of personal information included in the information extracted from the images captured by the camera. The information processing system according to claim 2.

4. The scope of disclosure is set to expand in stages, in the following order, as the range of personal information to be included in the information extracted from the image captured by the camera expands: skeletal information of a person obtained by estimating the skeleton of a person in the captured image; identification information of only persons who have given prior consent to the provision of personal information; and identification information of both persons who have given prior consent and persons who have not. The information processing system according to claim 2.

5. When an abnormal situation occurs and the information extracted from the captured image cannot be transferred to a pre-configured destination, the processor inputs the information extracted from the image captured by the camera into a large-scale language model that converts the input information into text information and outputs it, thereby obtaining text information about the image content of the captured image. The acquired text information is output as audio information via the audio output unit. The information processing system according to claim 1.

6. The processor performs skeletal estimation of a person in the captured image to detect whether or not there is a postural abnormality in the person. When an abnormal situation occurs, and information extracted from the captured image cannot be transferred to a pre-configured destination, and if a posture abnormality of a person is detected in the image, the information extracted from the image in which the posture abnormality of the person was detected is input to the large-scale language model to obtain text information regarding the state of the person in the captured image. The acquired text information is output as audio information via the audio output unit. The information processing system according to claim 5.

7. The process involves determining the scope of disclosure, which is determined based on the level of personal information protection required for cameras that continuously capture images of pre-set locations, and the urgency of any abnormal situations that may occur, and deciding to what extent personally identifiable information should be included in the information extracted from the captured images. Depending on the determined scope of disclosure, the process includes the step of transmitting information extracted from images captured by the camera, including personally identifiable information, to an external device, A program that causes a computer to execute something.