Handling verification of visual content

WO2026195185A1PCT designated stage Publication Date: 2026-09-24TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/066651
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-20
Filing Date
2025-06-13
Publication Date
2026-09-24

Smart Images

  • Figure EP2025066651_24092026_PF_FP_ABST
    Figure EP2025066651_24092026_PF_FP_ABST
Patent Text Reader

Abstract

There is provided a method for handling verification of visual content. The method is performed by a first node. The method comprises selecting (102), from a plurality of second nodes, a second node to verify a first object identified in a semantic representation of visual content. Selecting the second node is based on which second node of the plurality of second nodes is identified as the most capable of verifying the first object. The method comprises transmitting (104), to the selected second node, a part of the semantic representation and a request for the selected second node to verify whether the visual content comprises the first object. The part comprises the first object method for handling local model updates for a global machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] HANDLING VERIFICATION OF VISUAL CONTENT

[0002] Technical Field

[0003] The disclosure relates to methods for handling verification of visual content, and nodes configured to operate in accordance with those methods.

[0004] Background

[0005] There exist many scenarios in which visual object detection can be useful. One type of visual object detection involves displaying identified object types and their locations on a screen. This type of visual object detection is a well-established in the field and has widespread commercial applications, such as in self-driving cars.

[0006] However, solely relying on visual object detection can present challenges. In particular, false negatives can occur. For example, in a self-driving application, a false negative of a pedestrian may be detected where a self-driving car is approaching a bus that has a picture of a person on its rear, or a false negative of a red traffic light may be detected where a self-driving car is facing the sun that appears as a red circle in the distance. False negatives such as this in self-driving applications can be disruptive to the user experience and potentially dangerous. While it is possible to improve performance by retraining the visual object detector using identified false positives, this approach is time intensive and prone to errors due to the need for manual identification and labelling of the false positives.

[0007] A more promising alternative involves symbolically describing the scene and employing reasoning to interpret the events in the scene. In this respect, graph extraction from images has recently become popular. An example of such an approach uses Visual Question Answering (VQA), as descried in “Abduction of Domain Relationships from Data for VQA", by A. Chowdhury et al., International Conference on Logic Programming (ICLP) 2024. The approach involves the extraction of symbolic graphs from images and the subsequent augmentation of external sources of knowledge. The augmentation of external sources of knowledge allows for reasoning and planning types of use case, such as asking advanced questions about properties of objects in the scene. There thus exist approaches that involve the formalisation of knowledge extracted from visual content ingraphs and the subsequent processing of these graphs (e.g. using reasoning techniques). In other approaches, a neurosymbolic system can be used to predict future events by analysing video images. This may involve answering a user query in natural language (e.g. what shape is the Xthobject to collide with Object Y) by predicting dynamics of a physical system.

[0008] However, a key challenge with these approaches lies in grounding the models used to reliable ground truth data. This challenge becomes more complex in applications where ground truth originates from multiple sources. An example is in the validation of detected objects using ground truth data from various sources. In a use case drawn from telecommunications, visual object detection is applied to identified components at a radio base station. The ground truth data for validating these components originates from different vendors. For instance, a customer service provider (CSP) may supply a ground truth for site installation aspects (e.g. installation, power systems, etc.). An authorised service provider (ASP) may contribute a ground truth for the radio cabinet and tower, while a telecom equipment vendor may provide data for hardware components of the base station, including the baseband processor, router, and power supply.

[0009] Thus, existing approaches are unreliable and prone to errors, since visual object detection can misclassify objects. This can lead to potential issues, some of which can have serious consequences in certain applications. The existing approaches are also complex, inefficient, and resource intensive.

[0010] Summary

[0011] It is therefore an object of the disclosure to obviate or eliminate at least some of the above-described disadvantages associated with existing techniques.

[0012] Therefore, according to an aspect of the disclosure, there is provided a first method for handling verification of visual content. The first method is performed by a first node. The first method comprises selecting, from a plurality of second nodes, a second node to verify a first object identified in a semantic representation of visual content. Selecting the second node is based on which second node of the plurality of second nodes is identified as the most capable of verifying the first object. The first method comprises transmitting, to the selected second node, a part of the semantic representation and arequest for the selected second node to verify whether the visual content comprises the first object. The part comprises the first object.

[0013] According to another aspect of the disclosure, there is provided a second method for handling verification of visual content. The second method is performed by a second node. The second method comprises receiving a part of a semantic representation of visual content and a request for the second node to verify whether the visual content comprises a first object identified in the semantic representation. The part comprises the first object. The second node is selected from a plurality of second nodes based on which of the plurality of second nodes is identified as the most capable of verifying the first object. The second method comprises verifying whether the visual content comprises the first object based on a comparison of the part to a ground truth.

[0014] According to another aspect of the disclosure, there is provided a method performed by a system. The method performed by the system comprises the first method and the second method.

[0015] According to another aspect of the disclosure, there is provided a first node comprising processing circuitry configured to cause the first node to select, from a plurality of second nodes, a second node to verify a first object identified in a semantic representation of visual content. Selecting the second node is based on which second node of the plurality of second nodes is identified as the most capable of verifying the first object. The processing circuitry is configured to cause the first node to transmit, to the selected second node, a part of the semantic representation and a request forthe selected second node to verify whether the visual content comprises the first object. The part comprises the first object.

[0016] In some embodiments, the first node may comprise at least one memory for storing instructions which, when executed by the processing circuitry, cause the first node to operate according to the first method.

[0017] According to another aspect of the disclosure, there is provided a second node comprising processing circuitry configured to cause the second node to receive a part of a semantic representation of visual content and a request for the second node to verify whether the visual content comprises a first object identified in the semanticrepresentation. The part comprises the first object. The second node is selected from a plurality of second nodes based on which of the plurality of second nodes is identified as the most capable of verifying the first object. The processing circuitry is configured to cause the second node to verify whether the visual content comprises the first object based on a comparison of the part to a ground truth.

[0018] In some embodiments, the second node may comprise at least one memory for storing instructions which, when executed by the processing circuitry, cause the second node to operate in accordance with the second method.

[0019] According to another aspect of the disclosure, there is provided a system. The system comprises the first node and the second node.

[0020] According to another aspect of the disclosure, there is provided a computer program comprising instructions which, when executed by processing circuitry, cause the processing circuitry to perform one or both of the first method and the second method.

[0021] According to another aspect of the disclosure, there is provided a computer program product, embodied on a non-transitory machine-readable medium, comprising instructions which are executable by processing circuitry to cause the processing circuitry to perform one or both of the first method and the second method.

[0022] Thus, in the manner described above, techniques for handling verification of visual content are provided, which can improve the accuracy with which detected objects can be identified (e.g. classified). The techniques are more reliable and less prone to errors. Despite this, the techniques are efficient and require minimal resources.

[0023] Brief description of the drawings

[0024] For a better understanding of the techniques, and to show how they may be put into effect, reference will now be made, by way of example, to the accompanying drawings, in which:

[0025] Figure 1 is a block diagram illustrating a first node according to an embodiment;Figure 2 is a block diagram illustrating a first method performed by the first node according to an embodiment;

[0026] Figure 3 is a block diagram illustrating a second node according to an embodiment;

[0027] Figure 4 is a block diagram illustrating a second method performed by the second node according to an embodiment;

[0028] Figure 5 is a block diagram illustrating a system and a method performed by the system according to an embodiment;

[0029] Figure 6 is a schematic illustration of a semantic representation of visual content according to an embodiment;

[0030] Figure 7 is a schematic illustration of a first profile according to an embodiment;

[0031] Figure 8 is a schematic illustration of a second profile according to an embodiment; and

[0032] Figures 9-11 are signalling diagrams illustrating an exchange of signals in a system according to an embodiment.

[0033] Detailed Description

[0034] Generally, all terms used herein are to be interpreted according to their ordinary meaning in the relevant technical field, unless a different meaning is clearly given and / or is implied from the context in which it is used. All references to a / an / the element, apparatus, component, means, step, etc. are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any methods disclosed herein do not have to be performed in the exact order disclosed, unless a step is explicitly described as following or preceding another step and / or where it is implicit that a step must follow or precede another step. Any feature of any of the embodiments disclosed herein may be applied to any other embodiment, wherever appropriate. Likewise, any advantage of any of the embodiments may apply to any other embodiments, and vice versa. Other objectives, features and advantages of the enclosed embodiments will be apparent from thefollowing description.

[0035] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Other embodiments, however, are contained within the scope of the subject-matter disclosed herein, the disclosed subject-matter should not be construed as limited to only the embodiments set forth herein; rather, these embodiments are provided byway of example to convey the scope of the subject-matter to those skilled in the art.

[0036] In some instances, detailed descriptions of well-known methods, nodes, interfaces, circuits, and devices are omitted so as not to obscure the description with unnecessary detail. Those skilled in the art will appreciate that the functions described may be implemented in one or more nodes using hardware circuitry (e.g., analogue and / or discrete logic gates interconnected to perform a specialised function, Application Specific Integrated Circuits (ASICs), Programmable Logic Arrays (PLAs), etc.) and / or using software programs and data in conjunction with one or more digital microprocessors or general purpose computers. Nodes that communicate using an air interface also have suitable radio communications circuitry. Moreover, where appropriate the technology can additionally be considered to be embodied entirely within any form of computer-readable memory, such as solid-state memory, magnetic disk, or optical disk containing an appropriate set of computer instructions that would cause a processor to carry out the techniques described herein.

[0037] As mentioned earlier, there are described herein techniques for handling verification of visual content. The methods described herein can improve the accuracy with which detected objects can be identified (e.g. classified).

[0038] The techniques described herein involve one or both of a first node and a second node. A system can comprise one or both of the first node and the second node. The system can be a network. The network can be any type of network. For example, the network may be a communications or telecommunications network. The network can be a mobile network, such as a fifth generation (5G) mobile network, a sixth generation (6G) mobile network, or any other generation mobile network. The network can be a core network (e.g. a 5G core (5GC) network, or 6G core network) or a radio access network (RAN). The network can be a virtual network or an at least partially virtual network. The networkcan be terrestrial and / or non-terrestrial. Although some examples have been provided for the type of network, it will be understood that the network can be any other type of network.

[0039] Figure 1 illustrates a first node 10 in accordance with an embodiment. The first node 10 is for handling verification of visual content. The first node 10 referred to herein can refer to equipment capable, configured, arranged and / or operable to communicate directly or indirectly with any one or more of the second nodes referred to herein, and / or with other nodes or equipment to enable and / or to perform the functionality described herein. The first node 10 referred to herein can, for example, be a physical node (e.g. a physical machine) or a virtual node (e.g. a virtual machine (VM)). The first node 10 referred to herein may be implemented in a cloud environment. Herein, the first node 10 may be referred to as an “Explanation Router (ER)”.

[0040] As illustrated in Figure 1, the first node 10 comprises processing circuitry (or logic) 12. The processing circuitry 12 controls the operation of the first node 10 and can implement the method described herein in respect of the first node 10. The processing circuitry 12 can be configured or programmed to control the first node 10 in the manner described herein. The processing circuitry 12 can comprise one or more hardware components, such as one or more processors, one or more processing units, one or more multi-core processors and / or one or more modules. In particular implementations, each of the one or more hardware components can be configured to perform, or is for performing, individual or multiple steps of the method described herein in respect of the first node 10. The processing circuitry 12 can be configured to run software to perform the method described herein in respect of the first node 10. The software may be containerised according to some embodiments. Thus, the processing circuitry 12 may be configured to run a container to perform the method described herein in respect of the first node 10.

[0041] Briefly, the processing circuitry 12 of the first node 10 is configured to cause the first node 10 to select, from a plurality of second nodes 20, 30, 40, a second node 20 to verify a first object identified in a semantic representation of visual content. Selecting the second node 20 is based on which second node of the plurality of second nodes 20, 30, 40 is identified as the most capable of verifying the first object. The processing circuitry 12 of the first node 10 is configured to cause the first node 10 to transmit (e.g. via the communications interface 16 of the first node 10), to the selected second node 20, a partof the semantic representation and a request for the selected second node 20 to verify whether the visual content comprises the first object. The part comprises the first object.

[0042] As illustrated in Figure 1, the first node 10 may optionally comprise a memory 14. The memory 14 of the first node 10 can comprise a volatile memory ora non-volatile memory. The memory 14 of the first node 10 may comprise a non-transitory media. Examples of the memory 14 of the first node 10 include, but are not limited to, a random access memory (RAM), a read only memory (ROM), a mass storage media such as a hard disk, a removable storage media such as a compact disk (CD) or a digital versatile disk (DVD), and / or any other memory.

[0043] The processing circuitry 12 of the first node 10 can be communicatively coupled (e.g. connected) to the memory 14 of the first node 10. The memory 14 of the first node 10 may be for storing program code or instructions which, when executed by the processing circuitry 12 of the first node 10, cause the first node 10 to operate in the manner described herein in respect of the first node 10. For example, the memory 14 of the first node 10 may be configured to store program code or instructions that can be executed by the processing circuitry 12 of the first node 10 to cause the first node 10 to operate in accordance with the method described herein in respect of the first node 10. Alternatively or in addition, the memory 14 of the first node 10 can be configured to store any information, data, messages, requests, responses, indications, notifications, signals, or similar, that are described herein. The processing circuitry 12 of the first node 10 may be configured to control the memory 14 of the first node 10 to store any of the information, data, messages, requests, responses, indications, notifications, signals, or similar, that are described herein.

[0044] As illustrated in Figure 1, the first node 10 may optionally comprise a communications interface 16. The communications interface 16 of the first node 10 can be communicatively coupled (e.g. connected) to the processing circuitry 12 of the first node 10 and / or the memory 14 of the first node 10. The communications interface 16 of the first node 10 may be operable to allow the processing circuitry 12 of the first node 10 to communicate with the memory 14 of the first node 10 and / or vice versa. Similarly, the communications interface 16 of the first node 10 may be operable to allow the processing circuitry 12 of the first node 10 to communicate with any one or more nodes (e.g. any one or more of the second nodes) referred to herein and / or any other node. Thecommunications interface 16 of the first node 10 can be configured to transmit and / or receive any of the information, data, messages, requests, responses, indications, notifications, signals, or similar, that are described herein. The processing circuitry 12 of the first node 10 may be configured to control the communications interface 16 of the first node 10 to transmit and / or receive any of the information, data, messages, requests, responses, indications, notifications, signals, or similar, that are described herein.

[0045] Although the first node 10 is illustrated in Figure 1 as comprising a single memory 14, it will be appreciated that the first node 10 may comprise at least one memory (i.e. a single memory or a plurality of memories) 14 that operate in the manner described herein. Similarly, although the first node 10 is illustrated in Figure 1 as comprising a single communications interface 16, it will be appreciated that the first node 10 may comprise at least one communications interface (i.e. a single communications interface or a plurality of communications interfaces) 16 that operate in the manner described herein. It will also be appreciated that Figure 1 only shows the components required to illustrate an embodiment of the first node 10 and, in practical implementations, the first node 10 may comprise additional or alternative components to those shown.

[0046] Figure 2 illustrates a first method performed by a first node 10 in accordance with an embodiment. The first method is for handling verification of visual content. The first node 10 as described earlier with reference to Figure 1 can be configured to operate in accordance with the first method of Figure 2. The first method can be performed by or under the control of the processing circuitry 12 of the first node 10.

[0047] With reference to Figure 2, as illustrated by block 102, a second node 20 is selected from a plurality of second nodes 20, 30, 40. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may select the second node 20. The second node 20 is selected to verify a first object identified in a semantic representation of visual content. Selecting the second node 20 is based on which second node of the plurality of second nodes 20, 30, 40 is identified as the most capable of verifying the first object.

[0048] As illustrated by block 104 of Figure 2, a part of the semantic representation and a request is transmitted to the selected second node 20. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may transmit the part and the request to selected second node 20 (e.g. via the communications interface 16 of the firstnode 10). The request is for the selected second node 20 to verify whether the visual content comprises the first object. The part comprises the first object.

[0049] The selecting of the second node 20 based on which of the plurality of second nodes 20, 30, 40 is identified as the most capable of verifying the first object may comprise comparing a first profile of the visual content to a second profile of each second node of the plurality of second nodes 20, 30, 40, and selecting the second node 20 that has a second profile that has the most features in common with the first profile. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may compare the first profile to the second profile(s) and select the second node 20 in this way. The first profile may comprise information about the first object. The second profile of each second node 20, 30, 40 may comprise information about one or more objects that the second node is capable of verifying.

[0050] The second node 20 that has a second profile that has the most features in common with the first profile can be referred to as the second node 20 that has a second profile that most closely matches the first profile. The features can refer to one or both of exact features (e.g. objects) and types of features (e.g. types of object).

[0051] The selecting of the second node 20 that has a second profile that has the most features in common with the first profile may comprise selecting the second node 20 that has a second profile comprising information indicative that the second node 20 is capable of verifying the first object. The selecting of the second node 20 that has a second profile that has the most features in common with the first profile may comprise, in the absence of a second node 20 that has a second profile 800 comprising information indicative that the second node 20 is capable of verifying the first object, selecting the second node 20 that has a second profile comprising information indicative that the second node 20 is capable of verifying the same type of object as the first object.

[0052] The selecting of the second node 20 may comprise selecting the second node 20 that has the second profile that has the most features in common with the first profile (e.g. as described above), and that meets one or both of the following conditions: the second node 20 is located geographically closest to a location at which the visual content originated (which can be referred to a spatial condition), the second node 20 has a second profile updated at a time closest to a time of acquisition of the visual content(which can be referred to as a temporal condition), and the second node 20 has resources available to process the request (which can be referred to as a load-based condition). The second node 20 may have resources available to process the request if it has ample resources available to process the request. The resources can, for example, comprise compute resources, memory resources, transport resources, or any other resources, or any combination of these resources.

[0053] Although not illustrated in Figure 2, the method may comprise generating (e.g. extract) the first profile from the semantic representation of the visual content. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may generate the first profile.

[0054] The information about the first object may comprise information identifying the first object and one or both of a type of object that the first object is identified (e.g. classified) as and a relationship (e.g. a connection or association) of the first object to one or more other objects identified in visual content.

[0055] The selecting of the second node 20 based on which of the plurality of second nodes 20, 30, 40 is identified as the most capable of verifying the first object 612 may comprise selecting the second node 20 that is predicted, during a learning process (e.g. a reinforcement learning process, or deep reinforcement learning process), to maximise a reward for verifying the first object. The reward may be indicative of an extent to which the second node 20 verifying the first object aligns with a ground truth.

[0056] Although not illustrated in Figure 2, the method may comprise receiving, from a third node, a request to verify the first object and selecting the second node 20 in response to the request. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may receive the request from the third node (e.g. via the communications interface 16 of the first node 10). Herein, the third node may be referred to as an “Original Source (OS)”.

[0057] Although also not illustrated in Figure 2, the method may comprise receiving, from the selected second node 20, a validation result indicative of whether the visual content is verified as comprising the first object. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may receive the validation result from theselected second node 20 (e.g. via the communications interface 16 of the first node 10).

[0058] The method may comprise selecting, from the plurality of second nodes 20, 30, 40, two or more second nodes to verify the first object. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may select the two or more second nodes. The selecting of the two or more second nodes may be based on which two or more second nodes of the plurality of second nodes 20, 30, 40 is identified as the most capable of verifying the first object (e.g. as described earlier). The method may comprise receiving, from each of the two or more selected second nodes, a validation result indicative of whether the visual content is verified as comprising the first object. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may receive the validation results from the selected second nodes (e.g. via the communications interface 16 of the first node 10).

[0059] The method may comprise computing, from the received validation results, an overall validation result indicative of whether the visual content is verified as comprising the first object. More specifically, the first node 10 (e.g. the processing circuitry 12 of the first node 10) may compute the overall validation result. The overall validation result may be an average of the received validation results, a majority of the received validation results, or some other overall validation result.

[0060] Figure 3 illustrates a second node 20, 30, 40 in accordance with an embodiment. The second node 20, 30, 40 is for handling verification of visual content. The second node 20, 30, 40 referred to herein can refer to equipment capable, configured, arranged and / or operable to communicate directly or indirectly with the first node 10 referred to herein, and / or with other nodes or equipment to enable and / or to perform the functionality described herein. The second node 20, 30, 40 referred to herein can, for example, be a physical node (e.g. a physical machine) or a virtual node (e.g. a virtual machine (VM)). Herein, the second node 20, 30, 40 may be referred to as a “Verifier (VE)”.

[0061] As illustrated in Figure 3, the second node 20, 30, 40 comprises processing circuitry (or logic) 22. The processing circuitry 22 controls the operation of the second node 20, 30, 40 and can implement the method described herein in respect of the second node 20, 30, 40. The processing circuitry 22 can be configured or programmed to control the second node 20, 30, 40 in the manner described herein. The processing circuitry 22 cancomprise one or more hardware components, such as one or more processors, one or more processing units, one or more multi-core processors and / or one or more modules. In particular implementations, each of the one or more hardware components can be configured to perform, or is for performing, individual or multiple steps of the method described herein in respect of the second node 20, 30, 40. The processing circuitry 22 can be configured to run software to perform the method described herein in respect of the second node 20, 30, 40. The software may be containerised according to some embodiments. Thus, the processing circuitry 22 may be configured to run a container to perform the method described herein in respect of the second node 20, 30, 40.

[0062] Briefly, the processing circuitry 22 of the second node 20, 30, 40 is configured to cause the second node 20, 30, 40 to receive (e.g. via the communications interface 26 of the second node 20) a part of a semantic representation of visual content and a request for the second node 20, 30, 40 to verify whether the visual content comprises a first object identified in the semantic representation. The part comprises the first object. The second node 20, 30, 40 is selected from a plurality of second nodes based on which of the plurality of second nodes is identified as the most capable of verifying the first object. The processing circuitry 22 of the second node 20, 30, 40 is configured to cause the second node 20, 30, 40 to verify whether the visual content comprises the first object based on a comparison of the part to a ground truth.

[0063] As illustrated in Figure 3, the second node 20, 30, 40 may optionally comprise a memory 24. The memory 24 of the second node 20, 30, 40 can comprise a volatile memory or a non-volatile memory. The memory 24 of the second node 20, 30, 40 may comprise a non-transitory media. Examples of the memory 24 of the second node 20, 30, 40 include, but are not limited to, a random access memory (RAM), a read only memory (ROM), a mass storage media such as a hard disk, a removable storage media such as a compact disk (CD) or a digital versatile disk (DVD), and / or any other memory.

[0064] The processing circuitry 22 of the second node 20, 30, 40 can be communicatively coupled (e.g. connected) to the memory 24 of the second node 20, 30, 40. The memory 24 of the second node 20, 30, 40 may be for storing program code or instructions which, when executed by the processing circuitry 22 of the second node 20, 30, 40, cause the second node 20, 30, 40 to operate in the manner described herein in respect of the second node 20, 30, 40. For example, the memory 24 of the second node 20, 30, 40may be configured to store program code or instructions that can be executed by the processing circuitry 22 of the second node 20, 30, 40 to cause the second node 20, 30, 40 to operate in accordance with the method described herein in respect of the second node 20, 30, 40. Alternatively or in addition, the memory 24 of the second node 20, 30, 40 can be configured to store any information, data, messages, requests, responses, indications, notifications, signals, or similar, that are described herein. The processing circuitry 22 of the second node 20, 30, 40 may be configured to control the memory 24 of the second node 20, 30, 40 to store any of the information, data, messages, requests, responses, indications, notifications, signals, or similar, that are described herein.

[0065] As illustrated in Figure 3, the second node 20, 30, 40 may optionally comprise a communications interface 26. The communications interface 26 of the second node 20, 30, 40 can be communicatively coupled (e.g. connected) to the processing circuitry 22 of the second node 20, 30, 40 and / or the memory 24 of the second node 20, 30, 40. The communications interface 26 of the second node 20, 30, 40 may be operable to allow the processing circuitry 22 of the second node 20, 30, 40 to communicate with the memory 24 of the second node 20, 30, 40 and / or vice versa. Similarly, the communications interface 26 of the second node 20, 30, 40 may be operable to allow the processing circuitry 22 of the second node 20, 30, 40 to communicate with any one or more nodes (e.g. the first node 10) referred to herein and / or any other node. The communications interface 26 of the second node 20, 30, 40 can be configured to transmit and / or receive any of the information, data, messages, requests, responses, indications, notifications, signals, or similar, that are described herein. The processing circuitry 22 of the second node 20, 30, 40 may be configured to control the communications interface 26 of the second node 20, 30, 40 to transmit and / or receive any of the information, data, messages, requests, responses, indications, notifications, signals, or similar, that are described herein.

[0066] Although the second node 20, 30, 40 is illustrated in Figure 3 as comprising a single memory 24, it will be appreciated that the second node 20, 30, 40 may comprise at least one memory (i.e. a single memory or a plurality of memories) 24 that operate in the manner described herein. Similarly, although the second node 20, 30, 40 is illustrated in Figure 3 as comprising a single communications interface 26, it will be appreciated that the second node 20, 30, 40 may comprise at least one communications interface (i.e. a single communications interface or a plurality of communications interfaces) 26that operate in the manner described herein. It will also be appreciated that Figure 3 only shows the components required to illustrate an embodiment of the second node 20, 30, 40 and, in practical implementations, the second node 20, 30, 40 may comprise additional or alternative components to those shown.

[0067] Figure 4 illustrates a second method performed by a second node 20, 30, 40 in accordance with an embodiment. The second method is for handling verification of visual content. The second node 20, 30, 40 described earlier with reference to Figure 3 can be configured to operate in accordance with the second method of Figure 4. The second method can be performed by or under the control of the processing circuitry 22 of the second node 20, 30, 40 according to some embodiments.

[0068] With reference to Figure 4, as illustrated by block 202, a part of a semantic representation of visual content and a request for the second node is received. More specifically, the second node 20, 30, 40 (e.g. the processing circuitry 22 of the second node 20, 30, 40) may receive the part and the request (e.g. via the communications interface 26 of the second node 20, 30, 40). The request is to verify whether the visual content comprises a first object identified in the semantic representation. The part comprises the first object. The second node 20, 30, 40 is selected from a plurality of second nodes based on which of the plurality of second nodes is identified as the most capable of verifying the first object.

[0069] As illustrated by block 204 of Figure 4, it is verified whether the visual content comprises the first object based on a comparison of the part to a ground truth. More specifically, the second node 20, 30, 40 (e.g. the processing circuitry 22 of the second node 20, 30, 40) verifies whether the visual content comprises the first object.

[0070] The part may comprise information indicative of an identified relationship (e.g. an identified connection or association) that the first object has to one or more other objects identified in visual content. The ground truth may comprise information indicative of relationships between objects. The comparison of the part to the ground truth may comprise a comparison of the identified relationship to the information indicative of relationships between objects.

[0071] Although not illustrated in Figure 4, the method may comprise transmitting, to the firstnode 10, a validation result indicative of whether the visual content is verified as comprising the first object. More specifically, the second node 20 (e.g. the processing circuitry 22 of the second node 20) may transmit the validation result to the first node 10 (e.g. via the communications interface 26 of the second node 20).

[0072] The visual content referred to herein can comprise an image (e.g. one or more images), a video (e.g. one or more videos), a video frame (e.g. one or more video frames), any other visual content, or any combination of visual content.

[0073] The semantic representation referred to herein can be any representation of the visual content that represents the visual content semantically. The semantic representation can represent the identified objects and their relationships (e.g. connections or associations) in the visual content. The relationships can be with other identified objects in the visual content. The semantic representation can describe objects in the visual content and their relationships in a way that conveys a semantic meaning. The semantic representation may, for example, comprise a graph, a table, a hierarchical tree, a triple or triplets (e.g. using a resource description framework (RDF)), tensors, a matrix, or any other form of semantic representation.

[0074] There is also provided a system (e.g. network) comprising the first node 10 described herein and the second node 20, 30, 40 described herein. A method performed by the system comprises the method described herein in respect of the first node 10 and the method described herein in respect of the second node 20, 30, 40.

[0075] Figure 5 illustrates a system in accordance with an embodiment. The system illustrated in Figure 5 comprises a first node (“ER”) 10, and a plurality of second nodes 20, 30, 40 (“VE”). The system illustrated in Figure 5 may also comprise a third node (“OS”) 50.

[0076] As illustrated by arrow 502 of Figure 5, the third node 50 transmits a request to the first node 10. Thus, the first node 10 receives the request from the third node 50. The request is to verify a first object identified in a semantic representation (e.g. a graph) of visual content. The request may also be referred to herein as a validation request.

[0077] At step 504 of Figure 5, the first node 10 selects, from the plurality of second nodes 20, 30, 40, a second node 20 to verify the first object. Thus, the first node 10 may select thesecond node 20 in response to the validation request. The selection of the second node 20 is based on which second node of the plurality of second nodes 20, 30, 40 is identified as the most capable of verifying the first object (e.g. as described earlier).

[0078] As illustrated by arrow 506 of Figure 5, the first node 10 transmits, to the selected second node 20, a part of the semantic representation and a request for the selected second node 20 to verify whether the visual content comprises the first object. Thus, the selected second node 20 receives the part and the request. The part comprises the first object. Although only one second node 20 is represented as being selected in Figure 5 for simplicity, it will be understood that multiple second nodes may be selected.

[0079] The selected second node 20 verifies whether the visual content comprises the first object based on a comparison of the part to a ground truth.

[0080] As illustrated by arrow 508 of Figure 5, the selected second node 20 transmits, to the first node 10, a validation result indicative of whether the visual content is verified as comprising the first object. Thus, the first node 10 receives the validation result from the selected second node 20. In the case of two or more selected second nodes, each of the two or more selected second nodes may transmit their validation result to the first node 10 and thus the first node 10 may receive a validation result from each of the two or more selected second nodes. The validation result(s) can be referred to as feedback.

[0081] In the case of multiple validation results, at step 510 of Figure 5, the first node 10 computes, from the received validation results, an overall validation result indicative of whether the visual content is verified as comprising the first object. The overall validation result can, for example, be an average of the received validation results or a majority of the received validation result.

[0082] As illustrated by arrow 512 of Figure 5, the first node 10 transmits a response to the third node 50. Thus, the third node 50 receives the response from the first node 10. The response can comprise the validation result or the overall validation result.

[0083] Thus, it is possible to crowdsource the validation of the object detection to one or more (e.g. multiple) endpoints. The third node (“OS”) 50 can provide a request comprising visual content (e.g. a source image), a semantic representation (e.g. a detected objectsgraph) of the visual content, and the object(s) to be investigated. The third node 50 may extract the semantic representation (e.g. the detected objects graph) in any suitable manner, and a person skilled in the art will be aware of various techniques in this regard. The first node (“ER”) 10 may receive the request, profile the content, and find potential verifier candidates to provide an explanation of whether the object(s) to be investigated are properly identified. The second node(s) (“VE(s)”) 20, 30, 40 can provide their feedback to the first node 10, and the first node 10 may augment this feedback, providing an answer back to the third node 50.

[0084] Advantageously, the first node 10 selects one or more second nodes (e.g. out of a larger set of second nodes) to verify the semantic representation (e.g. the detected objects graph), and the one or more selected second nodes can compare part of that semantic representation (e.g. the detected objects graph) to a ground truth.

[0085] Figure 6 illustrates a semantic representation 600 of visual content 602 in accordance with an embodiment. The semantic representation is a first graph 600. The first graph 600 may be generated by the third node (“OS”) 50. For example, the third node 50 may extract the first graph 600 from the visual content 602. The visual content 602 may be a source image (as illustrated), or any other form of visual content. The first graph 600 may also be referred to herein as a detected objects graph. The first graph 600 can be sent as part of the validation request sent from the third node 50 to the first node 10.

[0086] The third node 50 can comprise a (visual) object detector for detecting the plurality of objects 604, 606, 608, 610, 612, 614, 616, 618 in the visual content 602. The object detector may identify (e.g. classify) the plurality of objects 604, 606, 608, 610, 612, 614, 616, 618 detected in the visual content 602. Thus, the first graph 600 comprises a plurality of objects 604, 606, 608, 610, 612, 614, 616, 618 identified in the visual content 602. The object detector may misidentify (e.g. misclassify) an object. As illustrated in Figure 6 by way of an example, the object detector has misidentified (e.g. misclassified) a first object 612 as a painting. The first object 612 is actually one of the windows.

[0087] The first graph 600 may comprise information indicative of relationships (e.g. connections or associations) between the plurality of objects 604, 606, 608, 610, 612, 614, 616, 618. As illustrated in Figure 6 by way of an example, the first graph 600 comprises information indicative that the house 606 contains a window 610, painting612, door 614, and roof 616, that a tree 608 is beside the house 608, and so on. The first graph 600 may comprise information indicative of a confidence with which an object has been identified. The information indicative of a confidence with which an object has been identified can be referred to herein as a confidence metric. The confidence metric can, for example, be a value (e.g. a percentage). As illustrated in Figure 6 by way of an example, the painting 612 has been identified with a confidence value of 0.73, whereas the window 610 has been identified with a confidence value of 0.94.

[0088] Figure 7 illustrates a first profile 700 of the visual content 602 in accordance with an embodiment. The first profile 700 may be generated by the first node (“ER”) 10. For example, the first node 10 may inject knowledge into the semantic representation 600 to build the first profile 700.

[0089] The first profile 700 comprises information about the plurality of objects to be verified. As illustrated in Figure 7 by way of an example, the first profile 700 may comprise information identifying the plurality of objects 708, 710, 712, 714 to be verified, and information indicative of the type of objects 720, 722, 724, 726 to be verified. The first profile 700 may comprise information indicative of relationships (e.g. connections or associations). As illustrated in Figure 7 by way of an example, the first profile 700 may comprise information indicative that the painting 712 is beside the window 710, that the door 714 is below the painting 712, that the house 708 contains the window 710, painting 712, and door 714, that the window 710 is a view component 724, that a building 722 has a view component 724, and so on.

[0090] Figure 8 illustrates a second profile 800 of a second node (“VE”) 20 in accordance with an embodiment. The second profile 800 may be generated by the second node 20. The second profile 800 can be referred to as a knowledge tree.

[0091] The second profile 800 illustrates the expertise of the second node 20. More specifically, the second profile 800 comprises information about a plurality of objects that the second node 20 is capable of verifying. As illustrated in Figure 8 by way of an example, the second profile 800 may comprise information identifying the plurality of objects 808, 812, 832, 834, 836 that the second node 20 is capable of verifying, and information indicative of the type of objects 820, 830, 824, 826 that the second node 20 is capable of verifying. The second profile 800 may comprise information indicative of relationships (e.g.connections or associations). As illustrated in Figure 8 by way of an example, the second profile 800 may comprise information indicative that a painting 812 is a type of art 830, that a house 808 is a type of structure 820, that a house 808 has a view component 824 and an access component 826, and so on.

[0092] Referring back to the examples illustrated in Figures 6 and 7, the semantic representation 600 indicated a relationship between the painting 612 and the house 608, and the first profile 700 also indicated a relationship between the painting 712 and the house 708. In contrast, the second profile 800 indicates no relationship between the painting 812 and the house 808. In particular, the second profile 800 indicates that the house 808 does not contain the painting 812. Thus, the second profile 800 can be used to verify that the painting 612, 712 has been misidentified (e.g. misclassified).

[0093] In addition to any of the information mentioned earlier, a profile (e.g. one or both of the first profile 700 and the second profile 800) may comprise any one or more of the following optional parameters:

[0094] An indication of an owner of the supplied information (e.g. an identification of the third node (“OS”) 50 that supplied the original source data and verifier-related information).

[0095] An indication of a location (e.g. a geographical location) of the owner. In terms of the second profile 800, the date the second profile 800 was last updated.

[0096] In terms of the first profile 700, the date the data for the first profile 700 was acquired.

[0097] An indication of the data provenance (e.g. an indication of who or which node supplied the information to the third node 50, and / or an indication of who or which node supplied the information to the second node 20).

[0098] In terms of the second profile 800, any geofencing rules that may prohibit any information contained in the second node 20 (e.g. stored at the second node 20) to be sent outside of one or more areas.

[0099] Figure 9 is a signalling diagram illustrating an exchange of signals in a system (e.g. network) according to an embodiment. The system illustrated in Figure 9 comprises a first node (“ER”) 10, and a second node 20 (“VE”). Although only one second node 20is illustrated in Figure 9, it will be understood that this is merely an example, and the system can comprise a plurality of second nodes. The system illustrated in Figure 9 may also comprise a third node (“OS”) 50. Although not illustrated in Figure 9, the system may also comprise a fourth node. Herein, the fourth node may be referred to as an “Ground Truth Original Source (GOS)”.

[0100] As illustrated by arrow 902 of Figure 9, the third node 50 transmits a request to the first node 10. Thus, the first node 10 receives the request from the third node 50. The request is to verify a first object 612 identified in a semantic representation 600 of visual content 602. The first object 612 may have been identified by the third node 50. The semantic representation 600 may have been generated by the third node 50. The request may comprise the semantic representation 600. The request can be referred to herein as a validation request. The validation request can trigger the process.

[0101] The third node 50 can identify (e.g. classify) objects. For example, the third node 50 may have a visual object detectorthat identifies (e.g. classifies) objects. The third node 50 (e.g. the visual object detector of the third node 50) may identify objects with a confidence threshold. The third node 50 (e.g. the visual object detector of the third node 50) can generate the semantic representation (e.g. a graph) 600 of the identified objects.

[0102] The validation request may be periodical. That is, the validation request may be transmitted periodically. In this example, it may be that random objects are chosen for validation. The validation request may be event-based. That is, the validation request may be transmitted when an event occurs. For example, the validation request may be transmitted when the third node 50 (e.g. the visual object detector of the third node 50) identifies an object with a confidence metric that is less than or equal to a threshold (e.g.

[0103] 75%). This is illustrated by means of an example in Figure 6, where the painting 612 is identified with a confidence metric that is less than the threshold. It may be that objects with a confidence metric that is less than or equal to the threshold are chosen for validation. Another trigger for the transmission of the validation request may be the variance in detected classes from the same source, which may indicate some type of adversarial attack.

[0104] Returning back to Figure 9, as illustrated by arrow 904, the first node 10 may generate (e.g. extract) a first profile 700 of the visual content 602 from the semantic representation600. The first profile 700 may thus be generated upon the first node 10 receiving the validation request. The first profile 700 is the profile of the semantic representation supplied by the third node 50. The first profile 700 is the profile that the first node 10 will compare with the second profile(s) of the second node(s) 20, 30, 40. An example of a first profile 600 is shown in Figure 7.

[0105] The first profile 700 can comprise information about the first object 712. The first profile 600 may comprise information about the type of content provided by the third node 50. The first profile 600 may start by identifying the objects that are to be investigated. These objects can either be determined at random, by use of a criterion of previous selection, or by use of a confidence metric from the third node 50 (e.g. the visual object detector of the third node 50).

[0106] Given the semantic representation (e.g. a graph of symbolically extracted relationships) from the third node 50 and a set of classes (objects) to be investigated, the first profile 700 may be built based on hierarchical abstraction. This approach may, for example, involve creating the first profile 700 by structuring objects within the semantic representation (e.g. graph) into a hierarchical framework, such as by using “isA” and “has” types of relationship. The hierarchical framework may associate a (generalised or abstract object) with its more specific or concrete counterparts.

[0107] An example of a first profile 700 is shown in Figure 7, where the dashed lines represent new knowledge added by the first node 10. The first node 10 may use internal and / or external vocabularies in order to inject this knowledge into the first profile 700. The first node 10 may create a more detailed or less detailed first profile 700, such as by exploring immediate relationships of the object or objects to be investigated (as shown in Figure 7) and optionally also n-th degree (e.g. 2nddegree) relationships.

[0108] Returning back to Figure 9, as illustrated by arrow 906, the first node 10 may acquire (e.g. receive or access) a second profile 800 of each second node of the plurality of second nodes 20, 30, 40. The first node 10 may maintain a list (e.g. an up-to-date list) of second profiles 800 of the plurality of second nodes 20, 30, 40. An example of a second profile 800 is shown in Figure 8.

[0109] The second profile 800 of each second node 20, 30, 40 can comprise information aboutone or more objects 812 that the second node is capable of verifying. The second profile 700 of each second node 20, 30, 40 can comprise information about the type of content handled by that second node. The second profile 800 can be similar to the first profile 700. A second profile 800 may be hardcoded to a second node (e.g. a priori). A second profile 800 can be discovered by the first node 10.

[0110] Returning back to Figure 9, as also illustrated by arrow 906, the first node 10 selects, from the plurality of second nodes 20, 30, 40, a second node 20 to verify the first object. Thus, the first node 10 may select the second node 20 in response to the request received from the third node 50. The selection of the second node 20 is based on which second node of the plurality of second nodes 20, 30, 40 is identified as the most capable of verifying the first object (e.g. as described earlier). For example, the first node 10 may compare the first profile 700 to (or with) the second profile 800 of each second node of the plurality of second nodes 20, 30, 40, and select the second node 20 that has a second profile 800 that has the most features in common with the first profile 700. As mentioned earlier, the second node 20 that has a second profile that has the most features in common with the first profile can be referred to as the second node 20 that has a second profile that most closely matches the first profile.

[0111] The first node 10 may select the second profile 800 (or the set of second profiles) that has the most features in common with, or that most closely matches, the first profile 700. It is the second node 20 attached to the selected profile 800 (or set of second nodes 20, 30, 40 attached to the set of second profiles) that are selected. The second profile 800 (or the set of second profiles) can be selected from a list of (candidate) second profiles.

[0112] The comparison of the first profile 700 to the second profile 800 of each second node of the plurality of second nodes 20, 30, 40 may be based on the degree of match within a hierarchy of those profiles. A match within a lower level in the hierarchy may indicate a closer overall match. For example, considering the first profile 700 illustrated in Figure 7 and the second profile illustrated in Figure 8, a match may be highly likely as the first object 712 to be investigated in the first profile 700 matches an object 812 in the second profile 800.

[0113] Additional factors may optionally be considered if a profile uses one or more of the aforementioned optional parameters. For example, if the first profile indicates that thedata provided originated from a first geographical location (e.g. the US) and a second profile indicates that a second node is located in a second geographical location (e.g. Europe) that is different from the first geographical location (e.g. a different country), or that is a threshold distance away from the first geographical location, then that second node may not be selected (e.g. due to geofencing issues). If a second profile is a close match to the first profile in terms of objects in the respective profiles, but the second profile has not been updated for an amount of time that exceeds a threshold amount of time (e.g. the second profile has not been updated for a long time), then that second profile may not be selected.

[0114] Thus, as illustrated by arrow 906 of Figure 9, the first node 10 selects a second node 20 to verify the first object. Although only one second node 20 is represented as being selected in Figure 9 for simplicity, it will be understood that multiple second nodes may be selected.

[0115] As illustrated by arrow 908 of Figure 9, the first node 10 transmits, to the selected second node 20, a part of the semantic representation and a request for the selected second node 20 to verify whether the visual content comprises the first object. Thus, the selected second node 20 receives the part and the request. The request can be referred to herein as a verification request. The part comprises the first object. Where there are multiple objects to verify, the part comprises each of those objects. The selected second node 20 verifies whether the visual content 602 comprises the first object based on a comparison of the part to a ground truth. In the case of two or more selected second nodes, each of the two or more selected second nodes may receive the part and the request from the first node 10, and verify whether the visual content 602 comprises the first object. The part may comprise the additional knowledge injected by the first node 10 that was mentioned earlier (e.g. to indicate relationships, such as hierarchical associations). Thus, the part of the semantic representation may be a part of the first profile 700.

[0116] The selected second node 20 may check whether its second profile 800 corresponds to reality (as represented in the received part). There are various different ways in which this check may be performed. For example, this check may involve a reasoning task, where the selected second node 20 validates each relationship of the first object 612, 712 against the second profile 800 of the selected second node. Alternatively or inaddition, the check may make use of a machine learning model (e.g. a neural network, such as a large language model (LLM)) to verify the outcome. The machine learning model may incorporate the second profile (which is a knowledge model) as an embedding, such as by using approaches for analysing semantic similarities (e.g. using the Graphs + Retrieval Augmented Generation (GraphRAG) approach, or similar).

[0117] It may be that the entire semantic representation is used instead of the part of the semantic representation. In an example where the semantic representation is a graph, the part of the semantic representation can be referred to as a subgraph.

[0118] Returning back to Figure 9, as illustrated by arrow 910, the selected second node 20 returns a response to the first node 10. More specifically, the second node 20 transmits, to the first node 10, a validation result indicative of whether the visual content is verified as comprising the first object. Thus, the first node 10 receives the validation result from the selected second node 20. In the case of two or more selected second nodes, each of the two or more selected second nodes may transmit their validation result to the first node 10 and thus the first node 10 may receive a validation result from each of the two or more selected second nodes.

[0119] As illustrated by arrow 912 of Figure 9, the first node 10 composes a response to the request received from the third node 50. The response can comprise the validation result received from the selected second node 20. In case multiple objects detected by the third node 50 are to be validated, the first node 10 may consolidate validations of multiple objects. Thus, in the case of multiple validation results, the first node 10 may compute, from the received validation results, an overall validation result indicative of whether the visual content is verified as comprising the first object. In this case, the response can comprise the overall validation result. The overall validation result can, for example, be an average of the received validation results (e.g. computed by means of simple averaging or weighted averaging) or a majority of the received validation result (e.g. computed using a majority voting scheme).

[0120] As illustrated by arrow 914 of Figure 9, the first node 10 transmits the response to the third node 50. Thus, the third node 50 receives the response from the first node 10.

[0121] Figure 10 is a signalling diagram illustrating an exchange of signals in a system (e.g.network) according to an embodiment. The system illustrated in Figure 10 comprises a first node (“ER”) 10, and a second node 20 (“VE”). Although only one second node 20 is illustrated in Figure 10, it will be understood that this is merely an example, and the system can comprise a plurality of second nodes. The system illustrated in Figure 10 may also comprise a third node (“OS”) 50. The system illustrated in Figure 10 may also comprise a fourth node (“GOS”) 60.

[0122] The exchange of signals in Figure 10 can be referred to as a training phase. The method performed by the system illustrated in Figure 10 uses learning (e.g. reinforcement learning, or deep reinforcement learning) at the first node 10 for learning which second node(s) to select.

[0123] In the method illustrated in Figure 10, second profiles 800 may not exist at all. Instead, after creating the first profile 700 at step 1004 of Figure 10, the first node 10 can learn which second profile 800 and thus which second node 20 to select. This learning occurs at steps 1006-1022. During steps 1006-1022, the first node 10 (or an agent within the first node 10) interacts with the environment, taking actions based on its current state and receiving rewards. Over time, the first node 10 (or an agent within the first node 10) learns to maximise rewards in any state.

[0124] In the context of Figure 10, actions may involve selecting second nodes. The selection can be based on the use case (e.g. the number of objects to verify, and / or whether multiple second nodes can be assigned per object). The state may correspond to the validation request or, more specifically, the object(s) to be verified. The validation request can be represented as a serialised (or vectorised) version of the semantic representation (e.g. graph) received from the third node 50 and the objects to be verified. The reward may reflect whether the selection aligns with a ground truth. The fourth node 60 is used during training to enable this. The fourth node 60 may be a node for receiving an input from a user (e.g. a human expert), a foundation machine learning model (e.g. LLM), or another system.

[0125] As illustrated by arrow 1002 of Figure 10, the fourth node 60 transmits a request to the first node 10. Thus, the first node 10 receives the request from the fourth node 60. The request is to verify a first object 612 identified in a semantic representation 600 of visual content 602. The first object 612 may have been identified by the fourth node 60. Thesemantic representation 600 may have been generated by the fourth node 60. The request may comprise the semantic representation 600. The request can be referred to herein as the validation request. An example of a graph (or objects graph) is listed as the semantic representation 600 in Figure 10. However, it will be understood that any other semantic representation is possible.

[0126] As illustrated by arrow 1004 of Figure 10, the first node 10 may generate (e.g. extract) a first profile 700 of the visual content 602 from the semantic representation 600. The first profile 700 can comprise information about the first object 712.

[0127] As illustrated by arrow 1006 of Figure 10, an action is performed whereby the first node 10 selects a second node 20. Although only one second node 20 is represented as being selected in Figure 10 for simplicity, it will be understood that multiple second nodes may be selected. The first node 10 may make the selection for the current state. The current state may be the validation request or, more specially, the first object 612, 712 that is to be verified.

[0128] In an initial iteration of the training phase, the method then proceeds to step 1014 of Figure 10. As illustrated by arrow 1014 of Figure 10, the first node 10 transmits, to the selected second node 20, a part of the semantic representation and a request for the selected second node 20 to verify whether the visual content comprises the first object 612, 712. Thus, the selected second node 20 receives the part and the request. In Figure 10, an example of a graph (or objects graph) is listed as the semantic representation 600, and a subgraph (or objects subgraph) is listed as the part of the semantic representation 600. However, it will be understood that any other semantic representation and part thereof is possible. The part comprises the first object 612, 712. The selected second node 20 verifies whether the visual content 602 comprises the first object based on a comparison of the part to a ground truth. In the case of two or more selected second nodes, each of the two or more selected second nodes may receive the part and the request from the first node 10, and verify whether the visual content 602 comprises the first object.

[0129] As illustrated by arrow 1016 of Figure 10, the selected second node 20 transmits, to the first node 10, a validation result indicative of whether the visual content is verified as comprising the first object. Thus, the first node 10 receives the validation result from theselected second node 20. In the case of two or more selected second nodes, each of the two or more selected second nodes may transmit their validation result to the first node 10 and thus the first node 10 may receive a validation result from each of the two or more selected second nodes.

[0130] As illustrated by arrow 1018 of Figure 10, the first node 10 composes a response to the request received from the third node 50. The response can comprise the validation result received from the selected second node 20. In the case of multiple validation results, the first node 10 may compute, from the received validation results, an overall validation result indicative of whether the visual content is verified as comprising the first object. In this case, the response can comprise the overall validation result. The overall validation result can, for example, be an average of the received validation results or a majority of the received validation result.

[0131] As illustrated by arrow 1020 of Figure 10, the first node 10 transmits the response to the fourth node 60. Thus, the fourth node 60 receives the response from the first node 10. The fourth node 60 may check the response against the ground truth to determine an extent to which the second node 20 verifying the first object aligns with the ground truth.

[0132] As illustrated by arrow 1022 of Figure 10, the fourth node 60 transmits a reward to the first node 10. Thus, the first node 10 receives the reward from the fourth node 60 for verifying the first object. The reward may be indicative of the extent to which the second node 20 verifying the first object aligns with the ground truth.

[0133] As illustrated by block 1024 of Figure 10, the current state becomes the old state, and it is this old state that is then stored at step 1008 of Figure 10 in the next iteration of the training phase.

[0134] In this next iteration of the training phase, steps 1002-1006 of Figure 10 are repeated. The first node 10 is in possession of a reward for the old state from step 1022 and, at step 1006 in this next iteration of the training phase, the first node 10 selects a second node 20 for the current state (e.g. the current validation request or, more specifically, the current object(s) to be verified). In each iteration of the training phase, at step 1006 of Figure 10, the first node 10 may select the second node 20 that is predicted to maximise the reward for verifying the current object(s), e.g. the first object 612, to be verified.As illustrated by arrow 1008 of Figure 10, the first node 10 may store an observation (or experience) in a memory, such as the memory 16 of the first node 10. For example, the first node 10 may store any one or more of the old state, the reward for the old state, the action in relation to the old state, the current state, and the action in relation to the current state. After the action in relation to the current state is performed at step 1006 to select a second node 20 for the current state, steps 1014-1022 may be performed in respect of the current state.

[0135] After a threshold number of (e.g. X) iterations of the training phase, as illustrated by arrow 1010 of Figure 10, the first node 10 may retrieve (e.g. pull) a plurality (e.g. M) of stored observations from the memory. Here, X may vary depending on the use case, and may be chosen depending on the complexity of the problem (e.g. X may be chosen to be higher in the case of higher resolution images, images with more objects, and / or more complicated semantic representations). An example forX is a value of 1000, but it will be understood that any other number of iterations is possible. As illustrated by arrow 1012 of Figure 10, the first node 10 may train a machine learning model (e.g. a neural network) using the retrieved observations. The machine learning model can be trained to select a second node 20 for verifying a given object.

[0136] For example, assume three second nodes (“VEs”) 20, 30, 40 with the following profiles (which are unknown to the first node (“ER”) 10):

[0137] VE1: [["industrial_building", "connected to", "warehouse"], ["shelves", "inside", "warehouse"], ["staking_robot", "inside", "industrial_building"], [conveyor_belt", "inside", "industrial_building"], ["shelves", "behind", "staking_robot"], ["conveyor_belt", "next to", "staking_robot"]]

[0138] VE2: [["Residential_Building", "contains", "Apartment"]. ["Apartment", "contains", "Window"], ["Apartment", "contains", "Door"], ["Apartment", "contains", "Couch"], ["Apartment", "contains", "Bed"], ["Apartment", "contains", "Kitchen"], ["Window", "part of', "Apartment"], ["Door", "part of', "Apartment"], ["Couch", "part of', "Apartment"], ["Bed", "part of', "Apartment"], ["Kitchen", "part of', "Apartment"]]

[0139] VE3: [[["Commercial_Building", "contains", "Store"], ["Store", "contains","Counter"], ["Store", "contains", "Point_of_Sale"], ["Counter", "part of', "Store"], ["Point_of_Sale", "part of', "Store"]]

[0140] Also, for a first iteration, assume the following input (“state”):

[0141] [(“box”, “on”, “conveyor_belt”), ("conveyor_belt", "next_to", "staking_robot"), ("shelves", "behind", "staking_robot")]

[0142] The action space, given the state is [VE1, VE2, VE3], Assume that the first node 10 selects VE2, in what is called an “action”.

[0143] To compare the profiles of the second nodes (VEs) 20, 30, 40 to the state, the first node 10 may calculate how much of the state is covered by the profile of each second node (VE) 20, 30, 40 using a matching metric. The state contains three relationships. VE1 covers two of them completely, specifically the relationships between "conveyor_belt" and "staking_robot," and "shelves" and "staking_robot". However, VE1 does not include the relationship “box on conveyor_belt”. This results in a normalised score of 0.67. VE2, however, does not provide any relevant information related to the entities or relationships in the state, as it is focused on residential apartments. Therefore, VE2 has a normalized score of 0.0. VE3 also does not contain any related entities or relationships (it focuses on commercial stores), so it too has a normalized score of 0.0. Thus, VE1 provides the most relevant information in this iteration with a score of 0.67, while VE2 and VE3 provide no relevant coverage of the state.

[0144] Thus, in this example, the first node 10 may record the following as an experience:

[0145] <state, action, reward, new_state> = <[(“box”, “on”, “conveyor_belt”), ("conveyor_belt", "next_to", "staking_robot"), ("shelves", "behind", "staking_robot")], VE2, 0, [“box”, “on”, “shelves”, ("conveyor_belt", "next_to", "staking_robot"), ("shelves", "behind", "staking_robot")]>

[0146] The bad reward of zero for selecting VE2 will incentivise the first node 10 to select another second node (verifier) if the same or similar state is encountered in the future. After a certain number (e.g. many) iterations of different variations of input, the first node 10 can implicitly learn to select the second node (verifier) that yields the best reward forthe same or similar state.

[0147] It may be that the training phase is iterated while the reward acquisition rate is greater than a threshold rate, or until a threshold number of (e.g. K) iterations of the training phase is reached. In an example where there are a total of 1000 iterations, then K may be a value of 50. However, it will be understood that other values of K are also possible.

[0148] Figure 11 is a signalling diagram illustrating an exchange of signals in a system (e.g. network) according to an embodiment. The system illustrated in Figure 11 comprises a first node (“ER”) 10, and a second node 20 (“VE”). Although only one second node 20 is illustrated in Figure 11, it will be understood that this is merely an example, and the system can comprise a plurality of second nodes. The system illustrated in Figure 11 may also comprise a third node (“OS”) 50. The system illustrated in Figure 11 may also comprise a fourth node (“GOS”) 60. The exchange of signals in Figure 11 can be referred to as an inference phase. The inference phase may follow the training phase of Figure 10, or may be separate from the training phase of Figure 10.

[0149] As illustrated by arrow 1102 of Figure 11, the third node 50 transmits a request to the first node 10. Thus, the first node 10 receives the request from the third node 50. The request is to verify a first object 612 identified in a semantic representation 600 of visual content 602. The first object 612 may have been identified by the third node 50. The semantic representation 600 may have been generated by the third node 50. The request may comprise the semantic representation 600. An example of a graph (or objects graph) is listed as the semantic representation 600 in Figure 11. However, it will be understood that any other semantic representation is possible.

[0150] As illustrated by arrow 1104 of Figure 11 , the first node 10 may generate (e.g. extract) a first profile 700 of the visual content 602 from the semantic representation 600. The first profile 700 can comprise information about the first object 712.

[0151] As illustrated by arrow 1106 of Figure 11 , the first node 10 selects, from the plurality of second nodes 20, 30, 40, a second node 20 to verify the first object. Thus, the first node 10 may select the second node 20 in response to the request received from the third node 50. The selection of the second node 20 is based on which second node of the plurality of second nodes 20, 30, 40 is identified as the most capable of verifying the firstobject (e.g. as described earlier). For example, the first node 10 may select the second node 20 that is predicted, during a learning process (e.g. the training phase in Figure 10), to maximise a reward for verifying the first object 612. The reward can be indicative of an extent to which the second node 20 verifying the first object 612 aligns with a ground truth. The first node 10 may make the selection for a given state. The given state may be the validation request. The first node 10 may make the selection using the machine learning model trained in the training phase.

[0152] Although only one second node 20 is represented as being selected in Figure 11 for simplicity, it will be understood that multiple second nodes may be selected.

[0153] As illustrated by arrow 1108 of Figure 11, the first node 10 transmits, to the selected second node 20, a part of the semantic representation 600 and a request for the selected second node 20 to verify whether the visual content comprises the first object. Thus, the selected second node 20 receives the part and the request. The part comprises the first object. In Figure 11 , an example of a graph (or objects graph) is listed as the semantic representation 600, and a subgraph (or objects subgraph) is listed as the part of the semantic representation 600. However, it will be understood that any other semantic representation and part thereof is possible. The selected second node 20 verifies whether the visual content 602 comprises the first object based on a comparison of the part to a ground truth. In the case of two or more selected second nodes, each of the two or more selected second nodes may receive the part and the request from the first node 10, and verify whether the visual content 602 comprises the first object.

[0154] As illustrated by arrow 1110 of Figure 11 , the selected second node 20 transmits, to the first node 10, a validation result indicative of whether the visual content is verified as comprising the first object. Thus, the first node 10 receives the validation result from the selected second node 20. In the case of two or more selected second nodes, each of the two or more selected second nodes may transmit their validation result to the first node 10 and thus the first node 10 may receive a validation result from each of the two or more selected second nodes.

[0155] As illustrated by arrow 1112 of Figure 11 , the first node 10 composes a response to the request received from the third node 50. The response can comprise the validation result received from the selected second node 20. In the case of multiple validation results,the first node 10 may compute, from the received validation results, an overall validation result indicative of whether the visual content is verified as comprising the first object. In this case, the response can comprise the overall validation result. The overall validation result can, for example, be an average of the received validation results or a majority of the received validation result.

[0156] As illustrated by arrow 1114 of Figure 11 , the first node 10 transmits the response to the third node 50. Thus, the third node 50 receives the response from the first node 10.

[0157] It may be the case that multiple second nodes 20, 30, 40 verify a single object. This can be the case, for example, when there is not a direct match between an object to be verified and the profiles of second nodes 20, 30, 40. In case multiple second nodes 20, 30, 40 match the object indirectly (e.g. by matching on some higher-level hierarchical object), then each of the second nodes 20, 30, 40 may be requested to provide verification results. As described earlier, these verification results can be used to compute an overall verification result (e.g. as an average of the verification results, or as a majority of the verification results).

[0158] There is also provided a computer program comprising instructions which, when executed by processing circuitry (such as the processing circuitry 12 of the first node 10 described herein and / or the processing circuitry 22 of the second node 20, 30, 40 described herein), cause the processing circuitry to perform at least part of the method described herein. There is provided a computer program product, embodied on a non-transitory machine-readable medium, comprising instructions which are executable by processing circuitry (such as the processing circuitry 12 of the first node 10 described herein and / orthe processing circuitry 22 of the second node 20, 30, 40 described herein) to cause the processing circuitry to perform at least part of the method described herein. There is provided a computer program product comprising a carrier containing instructions for causing processing circuitry (such as the processing circuitry 12 of the first node 10 described herein and / orthe processing circuitry 22 of the second node 20, 30, 40 described herein) to perform at least part of the method described herein. In some embodiments, the carrier can be any one of an electronic signal, an optical signal, an electromagnetic signal, an electrical signal, a radio signal, a microwave signal, or a computer-readable storage medium.In some embodiments, the first node functionality and / or second node functionality described herein can be performed by hardware. Thus, in some embodiments, the first node 10 and / or second node 20, 30, 40 described herein can be a hardware entity. However, it will also be understood that optionally at least part or all of the first node functionality and / or second node described herein can be virtualised. For example, the functions performed by the first node 10 and / or the second node 20, 30, 40 described herein can be implemented in software running on generic hardware that is configured to orchestrate the first node functionality and / or second node functionality described herein. Thus, in some embodiments, the first node 10 and / or the second node 20, 30, 40 described herein can be a virtual node. In some embodiments, at least part or all of the first node functionality and / or second node functionality described herein may be performed in a network enabled cloud. Thus, the method described herein can be realised as a cloud implementation according to some embodiments. The first node functionality and / or second node functionality described herein may all be at the same location or at least some of the first node functionality and / or second node functionality may be distributed, e.g. the first node functionality and / or second node functionality may be performed by one or more different nodes.

[0159] It will be understood that at least some or all of the method steps described herein can be automated in some embodiments. That is, in some embodiments, at least some or all of the method steps described herein can be performed automatically. The method described herein can be a computer-implemented method.

[0160] Therefore, as described herein, there are provided techniques for handling verification of visual content. The methods described herein can improve the accuracy with which detected objects can be identified (e.g. classified). The techniques described herein can provide efficient methods to integrate (e.g. symbolic) reasoning with reliable, multisource ground truth data (e.g. multi-source ground truth data) to address misclassifications in visual object detection without the high cost of manual labelling. The techniques described herein can be used for verifying the results of visual object detection against a ground truth without having to retrain a visual object detector. Also, the techniques described herein preserve the privacy between the second nodes that are used as verifiers, as a global source of knowledge is not necessary and, instead, multiple sources of ground truth that specialise in a specific knowledge set can be used.The techniques described herein can be applied to a variety of use cases. An example use case is self-driving cars, such as where visual object detection is key to the safe operation of the cars. Another use case is field service operations, such as where visual object detection may be key for engineers to detect equipment and fix it (e.g. radio base station maintenance). Another use case is in an industrial vertical within a smart factory. For example, an automated guided vehicle (AGV) may need to navigate a factory environment safely, and the techniques described herein can improve its ability to do so in an efficient manner. Although some use cases have been provided as examples, it will be understood that the techniques described herein are also applicable to a variety of other use cases in which objects need to be identified.

[0161] It should be noted that the above-mentioned embodiments illustrate rather than limit the idea, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. The word “comprising” does not exclude the presence of elements or steps other than those listed in a claim, “a” or “an” does not exclude a plurality, and a single processor or other unit may fulfil the functions of several units recited in the claims. Any reference signs in the claims shall not be construed so as to limit their scope.

Claims

36CLAIMS1. A method for handling verification of visual content, wherein the method is performed by a first node (10), the method comprising:selecting (102, 504, 906, 1106), from a plurality of second nodes (20, 30, 40), a second node (20) to verify a first object (612) identified in a semantic representation (600) of visual content (602), wherein selecting the second node (20) is based on which second node of the plurality of second nodes (20, 30, 40) is identified as the most capable of verifying the first object (612); andtransmitting (104, 506, 908, 1108), to the selected second node (20), a part of the semantic representation (600) and a request for the selected second node (20) to verify whether the visual content (602) comprises the first object (612), wherein the part comprises the first object (612).

2. The method as claimed in claim 1 , wherein:selecting (102, 504, 906) the second node (20) based on which of the plurality of second nodes (20, 30, 40) is identified as the most capable of verifying the first object (612) comprises:comparing a first profile (700) of the visual content (602) to a second profile (800) of each second node of the plurality of second nodes (20, 30, 40), wherein the first profile (700) comprises information about the first object (712) and the second profile (800) of each second node comprises information about one or more objects (812) that the second node is capable of verifying; and selecting the second node (20) that has a second profile (800) that has the most features in common with the first profile (700).

3. The method as claimed in claim 2, wherein:selecting the second node (20) that has a second profile (800) that has the most features in common with the first profile (700) comprises:selecting the second node (20) that has a second profile (800) comprising information indicative that the second node (20) is capable of verifying the first object (712).

4. The method as claimed in claim 2 or 3, wherein:selecting the second node (20) that has a second profile (800) that has the most features in common with the first profile (700) comprises:37in the absence of a second node (20) that has a second profile (800) comprising information indicative that the second node (20) is capable of verifying the first object (712), selecting the second node (20) that has a second profile (800) comprising information indicative that the second node (20) is capable of verifying the same type of object as the first object (712).

5. The method as claimed in any of claims 2 to 4, wherein:selecting (102, 504, 906) the second node (20) comprises:selecting the second node (20) that has the second profile (800) that has the most features in common with the first profile (700); andthat meets one or both of the following conditions:the second node (20) is located geographically closest to a location at which the visual content (602) originated;the second node (20) has a second profile (800) updated at a time closest to a time of acquisition of the visual content (602); and the second node (20) has resources available to process the request.

6. The method as claimed in any of claims 2 to 5, the method comprising:generating the first profile (700) from the semantic representation (600) of the visual content (602).

7. The method as claimed in any of claims 2 to 6, wherein:the information about the first object (712) comprises:the first object (712); andone or both of:a type of object that the first object (712) is classified as; and a relationship of the first object (712) to one or more other objects (708, 710, 714) identified in visual content (602).

8. The method as claimed in claim 1 , wherein:selecting (102, 504, 1106) the second node (20) based on which of the plurality of second nodes (20, 30, 40) is identified as the most capable of verifying the first object (612) comprises:selecting the second node (20) that is predicted, during a reinforcement learning process, to maximise a reward for verifying the first object (612).

9. The method as claimed in claim 8, wherein:the reward is indicative of an extent to which the second node (20) verifying the first object (612) aligns with a ground truth.

10. The method as claimed in any of the preceding claims, the method comprising:receiving (502, 902, 1102), from a third node (50), a request to verify the first object (612); andselecting (102, 504, 906, 1106) the second node (20) in response to the request.

11. A method as claimed in any of the preceding claims, the method comprising: receiving (508, 910, 1110), from the selected second node (20), a validation result indicative of whether the visual content (602) is verified as comprising the first object (612).

12. A method as claimed in claim 11 , the method comprising:selecting (102, 504, 906, 1106), from the plurality of second nodes (20, 30, 40), two or more second nodes to verify the first object (612), wherein selecting the two or more second nodes is based on which two or more second nodes of the plurality of second nodes (20, 30, 40) is identified as the most capable of verifying the first object (612); andreceiving (508, 910, 1110), from each of the two or more selected second nodes, a validation result indicative of whether the visual content (602) is verified as comprising the first object (612).

13. A method as claimed in claim 12, the method comprising:computing (510, 912, 1112), from the received validation results, an overall validation result indicative of whether the visual content (602) is verified as comprising the first object (612),wherein the overall validation result is an average of the received validation results or a majority of the received validation results.

14. The method as claimed in any of the preceding claims, wherein:the semantic representation comprises a graph, a table, a hierarchical tree, triplets, tensors, or a matrix.

15. A method for handling verification of visual content, wherein the method isperformed by a second node (20), the method comprising:receiving (506, 908, 1108) a part of a semantic representation (600) of visual content (602) and a request for the second node (20) to verify whether the visual content (602) comprises a first object (612) identified in the semantic representation (600), wherein the part comprises the first object (712) and the second node (20) is selected from a plurality of second nodes (20, 30, 40) based on which of the plurality of second nodes (20, 30, 40) is identified as the most capable of verifying the first object (612); andverifying whether the visual content (602) comprises the first object (612) based on a comparison of the part to a ground truth.

16. A method as claimed in claim 15, wherein:the part comprises information indicative of an identified relationship that the first object (712) has to one or more other objects (708, 710, 714) identified in visual content (602);the ground truth comprises information indicative of relationships between objects; andthe comparison of the part to the ground truth comprises a comparison of the identified relationship to the information indicative of relationships between objects.

17. A method as claimed in claim 15 or 16, the method comprising:transmitting (508, 910, 1110), to the first node (10), a validation result indicative of whether the visual content (602) is verified as comprising the first object (612).

18. The method as claimed in any claims 15 to 17, wherein:the semantic representation comprises a graph, a table, a hierarchical tree, triplets, tensors, or a matrix.

19. A method performed by a system, the method comprising:the method as claimed in any of claims 1 to 14; andthe method as claimed in any of claims 15 to 18.

20. A first node (10) comprising processing circuitry (12) configured to cause the first node (10) to:select, from a plurality of second nodes (20, 30, 40), a second node (20) to verify a first object (612) identified in a semantic representation (600) of visual content (602),wherein selecting the second node (20) is based on which second node of the plurality of second nodes (20, 30, 40) is identified as the most capable of verifying the first object (612); andtransmit, to the selected second node (20), a part of the semantic representation (600) and a request for the selected second node (20) to verify whether the visual content (602) comprises the first object (612), wherein the part comprises the first object (712).

21. A first node (10) as claimed in claim 20, wherein:the processing circuitry (12) is configured to cause the first node (10) to perform the method according to any of claims 2 to 14.

22. A second node (20) comprising processing circuitry (22) configured to cause the second node (20) to:receive a part of a semantic representation (600) of visual content (602) and a request for the second node (20) to verify whether the visual content (602) comprises a first object (612) identified in the semantic representation (600), wherein the part comprises the first object (712) and the second node (20) is selected from a plurality of second nodes (20, 30, 40) based on which of the plurality of second nodes (20, 30, 40) is identified as the most capable of verifying the first object (612); andverify whether the visual content (602) comprises the first object (612) based on a comparison of the part to a ground truth.

23. A second node (20) as claimed in claim 22, wherein:the processing circuitry (22) is configured to cause the second node (20) to perform the method according to any of claims 16 to 18.

24. A system comprising:the first node (10) as claimed in claim 20 or 21; andthe second node (20) as claimed in claim 22 or 23.

25. A computer program comprising instructions which, when executed by processing circuitry, cause the processing circuitry to perform the method according to any of claims 1 to 14, and / or any of claims 15 to 18.

26. A computer program product, embodied on a non-transitory machine-readable41medium, comprising instructions which are executable by processing circuitry to cause the processing circuitry to perform the method according to any of claims 1 to 14, and / or any of claims 15 to 18.