Machine vision system, machine vision method and machine vision apparatus
Patent Information
- Application Number
- TW114117731
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-05-12
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-05-11
AI Technical Summary
Existing machine vision systems infringe on personal privacy by requiring pre-stored facial images for identification and have limited visual range, leading to potential data leaks and restricted image recognition accuracy.
A machine vision system comprising multiple devices and a server that utilize privacy-preserving models for image analysis, de-identification, and federated learning to expand visual range and enhance recognition accuracy while protecting privacy.
The system achieves over 99% accuracy in real-time video monitoring with enhanced visual range and secure de-identification, reducing AI computing power consumption by 30% and ensuring privacy protection.
Smart Images

Figure TWG2TB001908826_001 
Figure TWG2TB001908826_002 
Figure TWG2TB001908826_003
Abstract
Description
[Technical Field]
[0001] This invention relates to a machine vision system, method, and apparatus. [Previous Technology]
[0002] Machine vision (MV) is an image processing technology widely used in industrial automation inspection, program control, and robot guidance. After acquiring monitoring images, a machine vision system can extract information from the images as needed. This information can be simple positive / negative information or complex datasets, such as the identity, position, and orientation of each object appearing in the image. In robot guidance applications, machine vision can integrate images from multiple cameras to automatically generate spatial information, enabling the robot to identify the position and orientation of objects in space, thereby performing tasks.
[0003] Existing machine vision systems identify targets in surveillance images by comparing them against a database. However, this method requires pre-storing facial images and identity data for retrieval or verification, severely infringing on personal privacy. Furthermore, if the stored data is leaked, personnel's identity information will be exposed, affecting their personal safety. In addition, existing machine vision systems can only construct visual information of their own surrounding space, limiting their visual range. Therefore, how to expand the visual range and improve image recognition accuracy while protecting personnel privacy is one of the important issues in this field. [Summary of the Invention]
[0004] The present invention provides a machine vision system, method and apparatus that can enhance visual recognition and task execution.
[0005] A machine vision system according to the present invention includes a plurality of machine vision devices and a server device. The plurality of machine vision devices are respectively configured to capture images of the spatial area where each machine vision device is located, and use a first machine learning model to analyze at least one target in the image and the correlation between each target and the spatial area. The server device receives the analysis results uploaded by each machine vision device and a plurality of first model parameters of the first machine learning model, and provides them to a second machine learning model to construct visual information of an overall space including all spatial areas. Each machine vision device downloads the visual information of the overall space and a set of second model parameters of the second machine learning model from the server device to update the first model parameters of the first machine learning model, and in response to a received task, generates instructions for executing the task using the updated first machine learning model.
[0006] In one embodiment of the present invention, the machine vision device includes using a first Privacy Visual Language Model (PVLM) to identify targets in an image and resolving the correlation between each target and a regional space to generate contextualized embeddings for each target in the regional space. The machine vision device further performs de-identification processing on the face images of each target to generate de-identified features, and compares the de-identified features with pre-stored features in a feature database to identify the identity of the target.
[0007] In one embodiment of the present invention, the above-mentioned machine vision device further uses a first privacy visual language model to analyze the human shape and actions of each target, and covers the human shape with a human shape mask to generate a de-identified image.
[0008] In one embodiment of the present invention, the machine vision device further embeds the regional context, action and identity of each target into the regional artificial intelligence model, and trains the regional artificial intelligence model using multiple tasks to generate a set of model parameters suitable for the regional artificial intelligence model to generate instructions to perform the task.
[0009] In one embodiment of the present invention, the above-mentioned regional context embedding includes the target's image token and text token, and generates image caption, image question answering, and space navigation between the target and the token.
[0010] In one embodiment of the present invention, the above-mentioned server device includes using a second privacy visual language model to fuse the parsing results uploaded by each machine vision device to generate multiple global context embeddings of each of the targets in the overall space.
[0011] In one embodiment of the present invention, the above-mentioned server device further uses the first model parameters of the first machine learning model uploaded by each machine vision device to train a global artificial intelligence model to generate the set of second model parameters suitable for recognizing all targets in the overall space.
[0012] In one embodiment of the present invention, the above-mentioned global artificial intelligence model includes federated learning using the first model parameters of the first machine learning model uploaded by each machine vision device to generate the set of second model parameters.
[0013] In one embodiment of the present invention, the above-mentioned machine vision devices are respectively configured in a plurality of corresponding user devices. Each machine vision device responds to the task received by the user device, captures the current image of the area space where the user device is located, uses the updated first machine learning model to analyze the targets in the current image of the area space and the correlation between each target and the area space, obtains the instructions for performing the task, and sends the instructions to the user device.
[0014] In one embodiment of the present invention, the above-mentioned machine vision devices are integrated with at least one of the corresponding user devices and server devices into a single device.
[0015] A machine vision method of the present invention is applicable to a machine vision system including multiple machine vision devices and a server device connected to each machine vision device. The method includes each machine vision device capturing an image of the area space where each machine vision device is located, using a first machine learning model to analyze at least one target in the image and the correlation between each target and the area space; the server device receiving the analysis results uploaded by each machine vision device and multiple first model parameters of the first machine learning model, and providing them to a second machine learning model to construct visual information of an overall space including all area spaces; and each machine vision device downloading the visual information of the overall space and a set of second model parameters of the second machine learning model from the server device to update the first model parameters of the first machine learning model, and generating instructions for performing the task using the updated first machine learning model in response to a received task.
[0016] In one embodiment of the present invention, the step of each machine vision device using a first machine learning model to analyze at least one target in an image and the correlation between each target and the regional space includes using a first privacy visual language model to identify targets in the image and analyzing the correlation between each target and the regional space to generate regional context embeddings of each target in the regional space. The step of each machine vision device using a first machine learning model to analyze at least one target in an image and the correlation between each target and the regional space further includes performing de-identification processing on the face images of each target to generate de-identification features and comparing the de-identification features with pre-stored features in a feature database to identify the identity of the target.
[0017] In one embodiment of the present invention, the step of each machine vision device using a first machine learning model to analyze at least one target in the image and the correlation between each target and the regional space further includes using a first privacy visual language model to analyze the human shape and action of each target, and covering the human shape with a human shape mask to generate a de-identified image.
[0018] In one embodiment of the present invention, the step of each machine vision device using a first machine learning model to analyze at least one target in an image and the correlation between each target and the regional space further includes embedding the regional context, action and identity of each target into a regional artificial intelligence model, and training the regional artificial intelligence model using multiple tasks to generate a set of model parameters suitable for the regional artificial intelligence model to generate instructions for performing tasks. The regional context embedding includes image tags and text tags of the target, and can perform visual language model applications such as image description, image question answering and spatial navigation between the target and the tags.
[0019] In one embodiment of the present invention, the step of constructing visual information of the overall space including all regional spaces by the server device using a second machine learning model includes using a second privacy visual language model to fuse the parsing results uploaded by each machine vision device to generate multiple global context embeddings of each target in the overall space.
[0020] In one embodiment of the present invention, the step of constructing visual information of the overall space including all regional spaces by the server device using the second machine learning model further includes training the global artificial intelligence model using the first model parameters of the first machine learning model uploaded by each machine vision device to generate the set of second model parameters suitable for recognizing all targets in the overall space.
[0021] In one embodiment of the present invention, the above-mentioned global artificial intelligence model includes federated learning using the first model parameters of the first machine learning model uploaded by each machine vision device to generate the set of second model parameters.
[0022] In one embodiment of the present invention, the above-mentioned machine vision devices are respectively configured in a plurality of corresponding user devices, and each machine vision device generates an instruction for performing the task using an updated first machine learning model in response to a received task. The steps include capturing the current image of the area space where the user device is located, using the updated first machine learning model to analyze the targets in the image of the area space and the correlation between each target and the area space, obtaining the instruction for performing the task, and sending the instruction to the user device.
[0023] A machine vision device of the present invention is configured in a user device and includes a communication device, a storage device, and a processor. The communication device is used to communicate with a server device. The storage device is used to store a plurality of first model parameters of a first machine learning model. The processor is coupled to the communication device and the storage device and is configured to capture images of the area space where the user device is located, use the first machine learning model to analyze at least one target in the image and the correlation between each target and the area space, and upload the analysis results to the server device. It also downloads visual information of the overall space including all area spaces and a set of second model parameters of a second machine learning model from the server device to update the plurality of first model parameters of the first machine learning model. The server device collects the analysis results uploaded by multiple machine vision devices and the first model parameters of the first machine learning model, and provides them to the second machine learning model to construct visual information of the overall space including all area spaces. In response to a task received by the user device, it generates instructions for performing the task using the updated first machine learning model and sends the instructions to the user device.
[0024] In one embodiment of the present invention, the above-described machine vision device is integrated with at least one of the user device and the server device into the same device.
[0025] Based on the above, the machine vision system, method, and apparatus of the present invention configure machine vision devices on user devices at the edge to capture and analyze images of the space where the user device is located, and the server device collects and integrates the analysis results of multiple machine vision devices to construct visual information of the overall space. Thus, by obtaining visual information of the overall space from the server device, the machine vision devices can enhance their visual recognition and task execution capabilities.
Implementation Method
[0027] The machine vision system provided in this embodiment of the invention is an innovative, privacy-conscious, multimodal plug-and-play intelligent robot system. It integrates technologies such as privacy-secure perception, multi-view fusion, cognitive inspiration, spatial intelligence, and robot learning. It uses federated learning, differential privacy, and homomorphic encryption technologies combined with artificial intelligence models to perform tasks, thereby protecting personal privacy and the security of sensitive information at the same time.
[0028] The machine vision device provided in this embodiment of the invention can be integrated with existing user devices such as cameras or robots through hardware interfaces such as Universal Serial Bus (USB) or Peripheral Component Interconnect Express (PCIe), providing comprehensive visual fusion and viewpoint coverage for user devices located at the edge. The multimodal privacy visual language model (PVLM) used in the machine vision device ensures efficient machine vision processing and secure face and human figure de-identification, while protecting the confidentiality of sensitive information and human privacy.
[0029] Figure 1A is a machine vision system architecture diagram according to an embodiment of the present invention, and Figure 1B is a schematic diagram of multi-view fusion according to an embodiment of the present invention. Referring to Figure 1A, the machine vision system 10 of the present invention collects views from various edge devices 14a-14j (including surveillance cameras, access control systems, robots, etc.) equipped with cameras or video cameras by the server device 12. Through online artificial intelligence (AI) learning, the edge devices 14a-14j can obtain more visual information, thereby effectively and accurately performing tasks. For example, the edge devices 14a-14j are configured at different locations on multiple floors of a building, and can capture views of different areas. The views captured by each edge device 14a-14j can be converted into contextualized embeddings through a privacy visual language model and shared with the server device 12.
[0030] Referring to Figure 1B, the server device 12, for example, employs multi-view fusion technology, using contextual embedding of multiple views from edge devices 14a-14g to reconstruct synthetic spatial vision through a visual base model, and generates visual information of the overall space, thereby expanding the visual range. For example, by redrawing the multiple views provided by edge devices 14a-14g, the scenes of each floor inside the building can be reconstructed. This visual information can be synchronously transmitted back to edge devices 14h-14j, enabling each edge device 14h-14j to accurately complete its assigned task using the expanded visual information.
[0031] FIG2 is a schematic diagram of a machine vision system according to an embodiment of the present invention. Referring to FIG2, the machine vision system 10 of the embodiment of the present invention includes a server device 12 and a plurality of machine vision devices 16.
[0032] The machine vision devices 16_1 to 16_n are connected to existing user devices 14_1 to 14_n equipped with cameras or camcorders via hardware interfaces such as Universal Serial Bus (USB) or PCIe, or integrated into the same device as the user devices 14_1 to 14_n. The user devices 14_1 to 14_n are, for example, edge devices such as IP cameras, access control systems, robotic vacuum cleaners, service robots, pet robots, and smart home appliances, or personal devices such as mobile phones, tablets, laptops, and desktop computers. This embodiment does not limit their type or quantity.
[0033] The server device 12 is, for example, a private server located in the cloud, which can collect visual information of the regional space (including regional context embedding, regional model parameters, etc.) uploaded by machine vision devices 16_1~16_n and perform fusion calculations to construct the visual information of the overall space and global model parameters.
[0034] Machine vision devices 16_1~16_n can download visual information of the entire space (including global context embedding, global model parameters, etc.) from server device 12 to provide comprehensive visual fusion and viewpoint coverage for user devices 14_1~14_n, thereby accurately completing the assigned tasks. In some embodiments, each machine vision device 16_1~16_n can be integrated with server device 12 into a single device to possess the function of collecting and fusing visual information provided by server device 12. That is, each machine vision device 16_1~16_n can operate independently even without server device 12, and can be connected in series with other machine vision devices 16_1~16_n to obtain visual information of the entire space.
[0035] In detail, FIG3 is a block diagram of a machine vision device according to an embodiment of the present invention. Referring to FIG3, the machine vision device 16 includes a communication device 162, a storage device 164, and a processor 166.
[0036] The communication device 162 includes, for example, a device that supports communication protocols such as wireless fidelity (Wi-Fi), radio frequency identification (RFID), Bluetooth, infrared, near-field communication (NFC), or device-to-device (D2D), or a device that supports Internet connectivity, for communicating with the server device 12. In some embodiments, the communication device 162 further includes a hardware interface such as Universal Serial Bus (USB) or High-Speed Peripheral Component Interconnect (PCIe) for connecting to or communicating with the user device 14 to acquire images captured by the user device 14.
[0037] The storage device 164 is, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk, or similar element or combination thereof, for storing computer programs executable by the processor 166. In some embodiments, the storage device 164 may also be used to store model parameters of machine learning models and a feature database recording pre-stored de-identification / encryption features (such as differential privacy, homomorphic encryption) of the target to be identified.
[0038] The processor 166 is, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose microprocessor, microcontroller, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), programmable logic device (PLD), or other similar device or combination of these devices. In some embodiments, the processor 166 may load a computer program from storage device 164 to execute the machine vision method of the embodiments of the present invention.
[0039] FIG4 is a flowchart illustrating a machine vision method according to an embodiment of the present invention. Please refer to FIG2, FIG3 and FIG4 simultaneously. The machine vision method of this embodiment is applicable to the machine vision system 10 of FIG2 and the machine vision device 16 of FIG3.
[0040] In step S402, each machine vision device 16 captures an image of the area space where the corresponding user device 14 is located, and uses the first machine learning model to analyze at least one target in the image and the correlation between each target and the area space.
[0041] In some embodiments, the first machine learning model includes a first privacy visual language model (PVLM). The processor 166 of the machine vision device 16 uses this first privacy visual language model to identify targets in an image and parse the correlation between each target and the regional space to generate contextualized embeddings for each target in the regional space. The contextualized embeddings include image tokens for the identified targets and text tokens to describe the targets. The first privacy visual language model can further link the identified people with objects in the space and perform scene parsing to determine which objects the person passes by and what actions they perform, i.e., to obtain the correlation between each target and the regional space.
[0042] In some embodiments, the first machine learning model further includes a regional artificial intelligence model. The processor 166 of the machine vision device 16 can input the regional artificial intelligence model by embedding the regional context of each target and the actions and identities of the identified targets, and train the regional artificial intelligence model using multiple tasks to generate a set of model parameters suitable for the regional artificial intelligence model to generate instructions for performing tasks.
[0043] In detail, FIG5 is a schematic diagram of artificial intelligence region image processing of a single machine vision device according to an embodiment of the present invention. Referring to FIG5, the processor 166 of the machine vision device 16, for example, captures an image of the region space where the corresponding user device 14 is located as a region view 51, and uses a privacy visual language model 52 to identify targets in the image and de-identify sensitive images (such as faces, human figures). The targets include objects such as people, tables, and chairs located in the region space. The privacy visual language model 52 generates image tags 54 and text tags 55 for each target by recognizing the contours, colors, sizes, and other features of each object in the image. The image tag 54 is the image of the target, and the text tag 55 is the text describing the target. The processor 166, for example, inputs the image tag YI and text tag YT of each target as a region context embedding (YI, YT) into the region artificial intelligence model 57, and trains the region artificial intelligence model 57 using various tasks 56 (including tasks 1 to n). Specifically, by embedding the regional context of each target and inputting the identity of the identified target into the regional artificial intelligence model 57, and training the regional artificial intelligence model using multiple tasks, model parameters 58 suitable for instructions to be executed by the regional artificial intelligence model are generated. By capturing images of the regional space and inputting them into the regional artificial intelligence model 57 during the execution of various tasks, the regional artificial intelligence model 57 can learn the actions (including instructions to execute the actions) required to perform the tasks in the regional space, thereby training the regional artificial intelligence model 57.
[0044] In detail, the Visual Language Model (VLM) achieves multimodal interaction and reasoning between text and images by fusing visual and linguistic information, and can be applied to various tasks such as image classification, text generation, image description, and spatial navigation. The privacy-preserving visual language model 52 in this embodiment adds a privacy protection mechanism to the traditional visual language model. When a person is identified from an image, de-identification processing is performed on the target's face image and / or human figure image, for example, by covering the human figure with a human figure mask to generate a de-identified image. Specifically, by converting the face image into de-identified features, the target's identity can be identified, and by converting the human figure image into de-identified features, the target's actions (including waving, standing, sitting, lying down, running, etc.) and their correlation with the regional space can be identified. Thus, while obtaining visual information of the required regional space, the privacy of people appearing in that regional space can be protected.
[0045] When the target is identified as a person, the processor 166 of the machine vision device 16 will further use the privacy visual language model 52 to perform de-identification processing on the face image of each target to generate de-identification features, and compare the de-identification features with the pre-stored features in the feature database stored in the storage device 164 to identify the identity of the target.
[0046] In some embodiments, the privacy visual language model 52 includes, for example, a deep learning (DL) model, which the processor 166 can use to perform de-identification processing on the region view 51. This deep learning model has object detection capabilities, can identify a target 53a in the input region view 51, and obscure the target 53a in the image to generate a de-identified image 53. Since the target 53a in the de-identified image has been obscured, even if the de-identified image 53 is leaked, a person viewing the de-identified image 53 will still be unable to identify the target 53a. Therefore, the de-identified image 53 can protect the privacy of the target 53a. In some embodiments, the deep learning model includes a deep neural network (DNN).
[0047] In some embodiments, the deep learning model can extract the face image of target 53a from the input image and perform a de-identification operation on the face image to generate one or more de-identification features. The processor 166, for example, uses an artificial intelligence model to determine whether the de-identification features match pre-stored features in a feature database to generate a verification result. The processor 166 may perform the de-identification operation based on, for example, a differential privacy algorithm to generate the de-identification features in a shorter time, or the processor 166 may perform the de-identification operation based on a homomorphic encryption algorithm or other encryption algorithms; no limitation is set herein. If the de-identification features match pre-stored features (e.g., the similarity between the de-identification features and pre-stored features is greater than a threshold), it means that target 53a's identity is a specific person corresponding to the pre-stored features. Accordingly, the processor 166 can generate a successful verification result. If the deidentified feature does not match any pre-stored feature (e.g., the similarity between the deidentified feature and the pre-stored feature is less than or equal to a threshold), it means that the identity of target 53a is unknown. Accordingly, processor 166 can generate a failed verification result. After generating the verification result, processor 166 can output the verification result for user reference.
[0048] To establish a feature database, processor 166 can acquire multiple historical images of multiple individuals and perform de-identification operations on the multiple historical images according to a deep learning model to generate multiple historical de-identified features. Processor 166 can establish a feature database based on the multiple historical de-identified features. The feature database may contain one or more historical de-identified features corresponding to the identity of a specific individual. The feature database is obtained, for example, by an embedded space or a loss function, such as AdaFace or ArcFace, which contains a margin for geodesic distance optimized by normalizing the correspondence of angles or radians in the normalized hypersphere.
[0049] On the other hand, the processor 166 may perform a de-identification operation on the face image of the target 53a to generate a de-identification tag, wherein the de-identification operation for generating the de-identification tag may be the same as or different from the de-identification operation for generating the de-identification feature, that is, the de-identification tag and the de-identification feature may be the same as or different. In some embodiments, the processor 166 may perform the de-identification operation for generating the de-identification tag based on, for example, a homomorphic encryption algorithm to generate a more easily identifiable de-identification tag, or the processor 166 may perform the de-identification operation based on other encryption algorithms (e.g., differential privacy algorithms). In one embodiment, the processor 166 may perform the de-identification operation based on a homomorphic encryption algorithm based on post-quantum-secure de-identification technology.
[0050] As shown in Figure 2, after the processor 166 of the machine vision device 16 completes the training of the regional artificial intelligence model, it can use the communication device 162 to upload the regional context embedding and the regional model parameters of the regional artificial intelligence model to the server device 12 through the privacy-security channel.
[0051] Returning to the process in Figure 4, in step S404, the server device 12 receives the parsing results uploaded by each machine vision device 16 and multiple first model parameters of the first machine learning model, and provides them to the second machine learning model to construct visual information of the overall space including all regional spaces.
[0052] In some embodiments, the server device 12, for example, uses a second privacy-preserving visual language model to fuse the parsing results uploaded by each machine vision device 16 to generate multiple global context embeddings for each target in the overall space. Furthermore, the server device 12 uses the first model parameters of the first machine learning models uploaded by each machine vision device 16 to train a global artificial intelligence model to construct complete visual information about the target. In some embodiments, the global artificial intelligence model, for example, uses the first model parameters of the first machine learning models uploaded by each machine vision device 16 to perform federated learning to generate a set of second model parameters. Alternatively, the global artificial intelligence model may average the first model parameters of the first machine learning models uploaded by each machine vision device 16 to generate a set of second model parameters.
[0053] In detail, FIG6 is a schematic diagram illustrating the training of a global artificial intelligence model according to an embodiment of the present invention. Referring to FIG6, after receiving the image tags 54, text tags 55 and model parameters 58 of the regional artificial intelligence model 57 uploaded by each machine vision device 16, the server device 12 uses the privacy visual language model 61 to fuse the image tags 54 and text tags 55 uploaded by each machine vision device 16 to construct the visual information 62 of the overall space and generate multiple global context embeddings of each target in the overall space, including image tags 63 and text tags 64.
[0054] On the other hand, the server device 12 also inputs the model parameters 58 of the received regional artificial intelligence model 57 into the global artificial intelligence model 65 to train the global artificial intelligence model 65 and generate global model parameters 66. The model parameters 58 of the regional artificial intelligence model 57 uploaded by each machine vision device 16 are the optimized parameters of the regional artificial intelligence model 57 after training, which include all the knowledge of the regional space. Therefore, after being trained by these model parameters 58, the global artificial intelligence model 65 possesses knowledge of the overall space including all regional spaces.
[0055] As shown in Figure 2, after the server device 12 completes the training of the global artificial intelligence model, it can provide the machine vision device 16 with the ability to download the global context embedding and the regional model parameters of the global artificial intelligence model through a privacy-security channel, so as to obtain visual information and knowledge of the overall space.
[0056] Returning to the flow in Figure 4, in step S406, each machine vision device 16 downloads visual information of the overall space and a set of second model parameters of the second machine learning model from the server device 12 to update the first model parameters of the first machine learning model. In response to the received task, each machine vision device 16 generates instructions for executing the task using the updated first machine learning model, but this is not limited to the above.
[0057] In some embodiments, the processor 166 of the machine vision device 16 updates the first privacy visual speech model, for example, using the downloaded visual information of the overall space (including global context embedding), and updates the model parameters of its own regional machine learning model using the model parameters of the downloaded global machine learning model.
[0058] Subsequently, the processor 166 of the machine vision device 16 responds to the task received by the user device 14. For example, it first captures the current image of the area where the user device 14 is located, uses the updated first privacy visual language model to analyze the target in the current image, and identifies the target's identity and actions. Specifically, the processor 166 performs de-identification processing on the face image of the target in the current image to generate de-identification features, and compares the de-identification features with pre-stored features in the feature database to identify the target's identity. In addition, the processor 166 can also perform action recognition on the human-shaped mask of the target in the current image to determine whether the target has dangerous actions, such as standing, sitting, running, etc. Combined with the analysis results (context embedding) of the first privacy visual language model, it can drive the regional artificial intelligence model to perform complex tasks and generate instructions for performing the tasks.
[0059] In detail, FIG7 is a schematic diagram illustrating the execution of a task according to an embodiment of the present invention. Referring to FIG7, when the robot 75 receives a task, the processor 166 of the machine vision device 16 captures the current image of the area space where the robot 75 is located, and uses an updated privacy visual language model to de-identify the captured image to obtain the de-identified image 71 and the correlation between the target 71a and the area space, and analyzes the actions 74 of each target 71a (including waving, lying down, standing, sitting, running or other specific actions), and then uses an updated area artificial intelligence model 73 to generate instructions to control the robot 75 to perform the task using the area context embedding, actions and identities of each target. Since the area artificial intelligence model 73 has learned the visual information 72 of the overall space, it can generate instructions suitable for the robot 75 to perform tasks in these area spaces based on the correlation between each target in the area space and the area space and the actions 74 of each target, and send these instructions to the robot 75 to control the robot 75 to perform the task according to the instructions. Furthermore, after updating the model parameters using the global artificial intelligence model, the regional artificial intelligence model 73 already possesses target information for all regions in the overall space, thus improving the accuracy of human / object identification.
[0060] Since the machine vision device 16 has acquired visual information of the entire space, its visual range has been expanded from the regional space to the entire space. Therefore, the types and scope of tasks it can perform can be extended to the entire space. The following are many application examples to illustrate the process of the machine vision device 16 performing tasks.
[0061] Task 1: When the manager enters the building and walks through the lobby towards the elevator, hand him the documents as he exits the elevator. Assuming the manager's office is on the 2nd floor, when the manager enters the building, the machine vision device installed in the lobby camera can identify the manager's identity and actions by analyzing the image captured by the lobby camera, and estimate the manager's walking and waiting time for the elevator. The analysis results are then uploaded to the cloud server. After collecting and integrating the analysis results uploaded by various machine vision devices, the cloud server provides the integrated visual information to the robot, enabling the robot to promptly obtain the documents and move to the elevator on the 2nd floor to wait, thus delivering the documents to the manager as he exits the elevator. The images uploaded to the cloud server have undergone facial and humanoid obfuscation processing, so even if someone else obtains the image, they cannot identify the person in it.
[0062] Task 2: When a customer sits down, deliver the menu to them; when a customer raises their hand, move to the customer's table to receive their order. Assuming a customer enters the restaurant, machine vision devices in multiple cameras deployed in the restaurant can identify each customer's actions by analyzing the images captured by the cameras and upload the analysis results to a cloud server. After collecting and integrating the analysis results uploaded by each machine vision device, the cloud server provides the integrated visual information to the robot. Therefore, when someone sits down in the restaurant, the robot, having obtained visual information, can determine the location of the customer who raised their hand and move to that location to deliver the menu. Similarly, when a customer raises their hand in the restaurant, the robot, having obtained visual information, can determine the location of the customer who raised their hand and move to that location to receive their order.
[0063] Task 3: Monitor the safety of someone entering the restroom. When someone enters the restroom, the machine vision device installed in the camera outside the restroom can identify the person's identity and actions by analyzing the images captured by the camera, and upload the analysis results to the cloud server. After collecting and integrating the analysis results uploaded by various machine vision devices, the cloud server can provide the integrated visual information to the cleaning robot and control the cleaning robot to switch to privacy mode to monitor the person's safety. For example, the cleaning robot will enter the restroom to check if the person has fallen or called for help after a predetermined time has passed. The images uploaded to the cloud server have undergone facial and human figure obfuscation processing, thereby protecting the privacy of the person entering the restroom.
[0064] In summary, the machine vision system, method, and apparatus of this invention employ a unique multimodal privacy visual language model (PVLM), which performs excellently in machine vision processing and secure de-identification, achieving an accuracy rate of over 99% in real-time video monitoring. The machine vision device can be easily integrated with existing devices or robots equipped with cameras or camcorders via hardware interfaces such as USB and PCIe, achieving comprehensive visual fusion, increasing image recognition accuracy to over 90%, and reducing overall AI computing power consumption by approximately 30%.
[0065] In machine vision devices, PVLM allows for real-time interaction with people and the environment, supporting precise machine intelligence tasks both online and offline. When interacting with robots, machine vision devices allow for the use of voice interfaces to prompt PVLM commands. Each device with a camera can contribute to unified visual understanding, thereby improving the overall accuracy of the system. By integrating advanced artificial intelligence and privacy technologies, machine vision devices provide powerful solutions for smart agencies and law enforcement, ensuring efficient task execution and robust privacy protection.
[0066] In the server-side device, federated learning and homomorphic encryption technologies ensure secure communication with the machine vision device. This setup allows for the secure merging and updating of context embeddings and model parameters, thereby enhancing visual recognition and task execution. This process ensures accurate event triggering and task completion while protecting sensitive information.
[0067] From the user's perspective, the machine vision device of this embodiment prioritizes privacy protection while providing efficient machine / robot intelligence capabilities through multi-view fusion. The multimodal PVLM model ensures that users can observe and track specific activities / behaviors (e.g., access control, security threat detection, or service request triggering) from multiple perspectives without compromising personal privacy. The user interface also allows authorized personnel to prompt robot / machine commands via voice input, making it a tool for users to easily interact with robots / machines.
[0068] From the perspective of materials and components, the machine vision device of this embodiment employs a powerful graphics processing unit (GPU), an optimized coupled multimodal deep neural network (DNN) and PVLM model, and multi-view fusion, enabling high-performance image processing and recognition tasks. Furthermore, the machine vision device utilizes privacy protection mechanisms, federated learning, differential privacy, and quantum-secure homomorphic encryption, which helps to minimize the ecological impact by reducing the risk of data leakage and unauthorized access.
[0069] The machine vision device is designed to operate on both edge and centralized / cloud computing platforms, with offline operation and online learning capabilities, thereby providing flexibility and scalability to meet diverse robotic service needs. The machine vision device can be easily configured via plug-and-play hardware and privacy-secure connections, enabling seamless updates and ensuring that edge devices stay up-to-date with the latest privacy-enhancing technologies.
[0070] Overall, the machine vision device of the present invention prioritizes user interests, privacy protection and ecological considerations, making it an advanced solution in the field of privacy-focused multimodal intelligent robot systems.
[0071] Based on the above, the machine vision system, method, and apparatus of the present invention can be applied to the following institutions / fields:
[0072] Law enforcement and security agencies: Used for surveillance, threat detection and access control, while ensuring privacy protection.
[0073] Healthcare: Data processing used in hospitals and clinics to monitor patient activity and ensure security.
[0074] Smart Home / Office: Enhance productivity, security, automation, and environmental monitoring in residential and office spaces.
[0075] Smart City: Used for traffic management, public safety and environmental monitoring.
[0076] Retail and shopping malls: Enhance security and customer experience through smart monitoring and service automation.
[0077] Manufacturing and warehousing: Improve operational efficiency and safety through robotic assistance and real-time monitoring.
[0078] Educational institutions: Used for campus security and smart infrastructure management.
[0079] Transportation: Safety and operational management at airports, railway stations and ports.
[0080] Government agencies: Used for secure data processing and monitoring of public spaces, while also protecting privacy. [Simplified Explanation of the Diagram]
[0026] Figure 1A is a machine vision system architecture diagram according to an embodiment of the present invention. Figure 1B is a schematic diagram of multi-view fusion according to an embodiment of the present invention. Figure 2 is a schematic diagram of a machine vision system according to an embodiment of the present invention. Figure 3 is a block diagram of a machine vision device according to an embodiment of the present invention. Figure 4 is a flowchart of a machine vision method according to an embodiment of the present invention. Figure 5 is a schematic diagram of training a regional artificial intelligence model according to an embodiment of the present invention. Figure 6 is a schematic diagram of training a global artificial intelligence model according to an embodiment of the present invention. Figure 7 is a schematic diagram of task execution according to an embodiment of the present invention.
Claims
1. A machine vision system, comprising: Multiple machine vision devices are configured to capture images of the area space where each machine vision device is located, and a first machine learning model is used to analyze at least one target in the image and the correlation between each target and the area space. The system includes a server device that receives the parsing results uploaded by each of the machine vision devices and multiple first model parameters of the first machine learning model, and provides them to a second machine learning model to construct visual information of an overall space including all regional spaces. Each of the machine vision devices downloads the visual information of the overall space and a set of second model parameters of the second machine learning model from the server device to update the first model parameters of the first machine learning model. In response to a received task, the machine vision device generates instructions to execute the task using the updated first machine learning model. The machine vision device includes using a first privacy visual language model (PVLM) to identify the targets in the image and parsing the correlation between each target and the regional space to generate regional contextualized embeddings of each target in the regional space. It also performs de-identification processing on the face images of each target to generate de-identified features and compares the de-identified features with pre-stored features in a feature database to identify the identity of the target.
2. The machine vision system of claim 1, wherein the machine vision device further uses the first privacy visual language model to analyze the human form and actions of each of the targets, and covers the human form with a human form mask to generate a de-identified image.
3. The machine vision system of claim 2, wherein the machine vision device further inputs the region context embedding, the action, and the identity of each of the targets to a training region artificial intelligence model, and trains the region artificial intelligence model using multiple tasks to generate a set of model parameters suitable for the instructions to be performed by the region artificial intelligence model on the tasks.
4. The machine vision system as claimed in claim 1, wherein the region context embedding includes an image token and a text token of the target, and the first privacy-preserving visual language model further generates an image caption, image question answering, and space navigation between the target and the image token or the text token.
5. The machine vision system as claimed in claim 1, wherein the server device includes fusing the parsing results uploaded by each of the machine vision devices using a second privacy visual language model to generate multiple global context embeddings of each of the targets in the overall space.
6. The machine vision system of claim 5, wherein the server device further uses the first model parameters of the first machine learning model uploaded by each of the machine vision devices to train a global artificial intelligence model to generate the set of second model parameters suitable for recognizing all targets in the overall space.
7. The machine vision system of claim 6, wherein the global artificial intelligence model includes federated learning using the first model parameters of the first machine learning model uploaded by each of the machine vision devices to generate the set of second model parameters.
8. The machine vision system as claimed in claim 1, wherein the machine vision devices are respectively configured in a plurality of corresponding user devices, each of the machine vision devices, in response to the task received by the user device, captures the current image of the area space where the user device is located, uses an updated first machine learning model to analyze the target in the current image of the area space and the correlation between each target and the area space, obtains the instruction for performing the task, and sends the instruction to the user device.
9. The machine vision system as claimed in claim 8, wherein each of the machine vision devices is integrated with at least one of the corresponding user device and server device into a single device.
10. A machine vision method applicable to a machine vision system comprising a plurality of machine vision devices and a server device connected to each of the machine vision devices, the method comprising: The process involves: each machine vision device capturing images of its respective region; using a first machine learning model to analyze at least one target in the image and the correlation between the target and the region; including: using a first privacy-preserving visual language model to identify the target in the image and analyzing the correlation between the target and the region to generate a regional context embedding for each target in the region; performing de-identification processing on the face images of each target to generate de-identified features, and comparing the de-identified features with pre-stored features in a feature database to identify the target's identity; receiving the analysis results uploaded by each machine vision device and multiple first model parameters of the first machine learning model from the server device, and providing them to a second machine learning model to construct visual information of an overall space including all region spaces; and each machine vision device downloading the visual information of the overall space and a set of second model parameters of the second machine learning model from the server device to update the first model parameters of the first machine learning model, and generating instructions to execute the task using the updated first machine learning model in response to a received task.
11. The method of claim 10, wherein the step of using a first machine vision device to analyze at least one target in the image and the spatial association between each target and the region further comprises: The first privacy visual language model is used to analyze the human form and actions of each of the targets, and a human form mask is applied to the human form to generate a de-identified image.
12. The method of claim 11, wherein the step of using a first machine vision device to analyze at least one target in the image and the spatial association between each target and the region further comprises: The region context embedding, the action, and the identity of each of the targets are input into a training region AI model, and the region AI model is trained using multiple tasks to generate a set of model parameters suitable for the instructions to be executed by the region AI model, wherein the region context embedding includes image tags and text tags of the targets, and a visual language model application is performed including at least one of image description, image question answering, and spatial navigation between the targets and the image tags or the text tags.
13. The method of claim 10, wherein the step of constructing visual information of the overall space including all regional spaces by the server device using the second machine learning model includes: The parsing results uploaded by each of the machine vision devices are fused using a second privacy visual language model to generate multiple global context embeddings for each of the targets in the overall space.
14. The method of claim 13, wherein the step of constructing visual information of the overall space including all regional spaces by the server device using the second machine learning model further comprises: The first model parameters of the first machine learning model uploaded by each of the machine vision devices are used to train a global artificial intelligence model to generate the set of second model parameters suitable for recognizing all targets in the overall space.
15. The method of claim 14, wherein the global artificial intelligence model includes federated learning using the first model parameters of the first machine learning model uploaded by each of the machine vision devices to generate the set of second model parameters.
16. The method of claim 10, wherein the machine vision devices are respectively configured in a plurality of corresponding user devices, and each of the machine vision devices, in response to a received task, generates instructions for performing the task using an updated first machine learning model, comprising: The system captures the current image of the area space where the user device is located, uses an updated first machine learning model to analyze the targets in the current image of the area space and the correlation between each target and the area space, obtains the instructions for performing the task, and sends the instructions to the user device.
17. A machine vision device, configured in a user device, comprising: A communication device that communicates with the server device. Storage device for storing multiple first model parameters of the first machine learning model; The system includes a processor coupled to the storage device and configured to: capture an image of the area space where the user device is located; use a first machine learning model to analyze at least one target in the image and the correlation between each target and the area space; and upload the analysis results to a server device; download visual information of the overall space including all area spaces and a set of second model parameters of a second machine learning model from the server device to update multiple first model parameters of the first machine learning model, wherein the server device collects the analysis results uploaded by multiple machine vision devices and the first model parameters of the first machine learning model, and provides them to the second machine learning model to construct the visual information of the overall space including all area spaces; In response to a task received by the user device, the processor generates instructions for performing the task using an updated first machine learning model, sends the instructions to the user device, wherein the processor includes using a first privacy visual language model to identify the targets in the image, and resolving the correlation between each target and the regional space to generate a regional context embedding for each target in the regional space, performing de-identification processing on the face images of each target to generate de-identified features, and comparing the de-identified features with pre-stored features in a feature database to identify the identity of the target.
18. The machine vision device as claimed in claim 17, wherein the machine vision device is integrated with at least one of the user device and the server device into the same device.
Citation Information
Patent Citations
Detection system for detecting falling of person in living space and detection method thereof
CN115331283A
Logistics personnel identity recognition method and device based on privacy protection and related medium
CN115527252A
Privacy protection face recognition system based on federated learning and frequency domain transformation
CN118053210A
Bystander-centric privacy controls for recording devices
TW202305661A
Privacy protection of digital image data on a social network
US20240119184A1