Cognitive reinforcement and verification of internet borne remote events in visual data
The system addresses the lack of agency in news consumption by enabling 3D navigation and direct communication with sources, allowing users to explore and verify news stories through a 3D space and video access.
Patent Information
- Application Number
- PCT/US2025/023975
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-04-09
- Publication Date
- 2025-10-16
AI Technical Summary
Existing technologies lack effective methods for verifying and authenticating remote events in visual data, particularly in the news industry, where consumers are reliant on media entities to decide what they see and do not see, lacking agency in exploring news sources.
A system and methodology for constructing a 3D space from multiple perspectives, integrating icons for accessing related information, enabling users to navigate and 'drill down' to source videos, and allowing one-on-one communication with sources for enhanced context and verification.
Empowers users with agency to explore and verify news stories by accessing multiple video streams and engaging directly with sources, enhancing understanding and context of events.
Smart Images

Figure US2025023975_16102025_PF_FP_ABST
Abstract
Description
[0001] PCT APPLICATION
[0002] A.LRajasingham. 6024 Bradley Boulevard, Bethesda MD 20817. air@mmmmg.com
[0003] TITLE OF IN ENTION: Cognitive reinforcement and verification of Intenet bourne remote events in visual data.
[0004] PRIORITY CLAIMED REFERENCES: This application claims priority from and incorporates by reference in their entirety the following provisional applications. 63 631 673 filed 2024-04-09; 63 652 313 filed 2024-05-28; 63 656 143 filed 2024-06-05
[0005] BACKGROUND OF INVENTION
[0006] FIELD OF INVENTION
[0007] Methodologies for virtual visualization and navigation in a remote physical space and provides a mechanism for verification and authentication of video and other evidence.
[0008] SUMMARY
[0009] The Drawings illustrate embodiments of the inventions. These features and more are described below. The invention relates to the referenced filed applications.
[0010] BRIEF DESCRIPTION OF DRAWINGS
[0011] Fig 27-01 illustrates the recording of the scene from the first perspective.
[0012] Fig 27-02 illustrates the recording of the scene from the second perspective
[0013] Fig 27-03 illustrates the recording of the scene from the third perspective
[0014] Fig 27-04 illustrates a first of multiple 3D views navigable in the available 3D rendering, showing the icons for selection of available recorded videos from the marked locations.
[0015] Fig 27-05 illustrates a second of multiple 3D views navigable in the available 3D rendering, showing the icons for selection of available recorded videos from the marked locations.
[0016] Fig 27-06 illustrates a third of multiple 3D views navigable in the available 3D rendering, showing the icons for selection of available recorded videos from the marked locations.
[0017] Fig l- l represents the replay of video when icon 1, is selected on any of the 3D views available.
[0018] Fig 27-08 represents the replay of video when icon 2, is selected on any of the 3D views available.
[0019] Fig 27-09 represents the replay of video when icon 2, is selected on any of the 3D views available.
[0020] Notably, the extent or boundaries of the representation of the 3D space will depend on the cumulative cover of all the recording views available.
[0021] Fig 27-10 Map of local region
[0022] Fig 27-11 - Composite view with map and video or image
[0023] Fig 27-12, represents a map with multiple source users with their sources randomly distributed around the space. This is essentially a population that is endowed with the network on their interface which can be a mobile phone. There is an optional viewer / user that has access to this map.
[0024] Fig 27-13, represents a map with the same multiple source users with their sources as in the previous figure, where one of the source users have seen an event of interest ("first source"), and has switched on his or her camera. (In an alternative embodiment there may be other machine based source user such as a drone or a surveillance camera, that can go to the same processes)
[0025] In an embodiment of the invention, as shown in figure 27-14, a predetermined number of sources with the closest Geo locations to the First Source ("second sources") are notified to switch on their cameras. (Such notification may be accompanied in some embodiments by an automatic switch on at their cameras.) In some such embodiments, the second sources see one or both of: the video from the camera of the first source, thereby understanding the event that needs to be covered; and a local map with the first source highlighted and their own position highlighted among the other sources that have their cameras on. This enables them to understand how to get to the event of interest.
[0026] Fig 27-15, shows the interaction with the viewer / user, where the viewer / user can have two way communication with any one or more of the first source and the second sources, while benefiting from the recording of these source streams.
[0027] Fig 27-16, 17 are of a presentation to a News consumer with a dynamic map of sources and the views of each source and a view of the News anchor or News room or journalist collating source information to build a coherent story. Each of these elements of visual data are represented on dynamic thumb nails (static for images). One of the views or the map are depicted on the main screen for careful examination. Selection of any of the thumb nails will put that thumb nail on the main screen and revert the current screen image to a thumb nail. It is also possible in some embodiments to have all the views permanently on the thumb nails and one of them by selection displayed on the main screen. Here Fig 27-16 depicts a dynamic map where the sources can move are tracked with the icons. Selection of the icons put the selected source view on the main screen. Alternatively, the selection of a thumb-nail puts the thumb nail video on the main screen (Image if static)
[0028] During the process of construction of the story by the News Anchor / Journalist, he / she has access to each of the sources and can have one on one conversations with each one of them by selecting them on the map or using the thumb nail. In some embodiments, the other Sources are muted to avoid excessive noise. The News Anchor / Journalist has voice channels to all the sources concurrently otherwise to coordinate the project.
[0029] Fig 27-18 is a screen of a video of an anchor / coordinating journalist that may be posted on social media or on portals of media outlets. While the story will play in the video on the main screen, it has a unique feature of an icon "On the Spot" ss seen on the bottom right-hand side, that when selected will provide the perspective similar to what is seen in Fig 27-17, with the anchor / coordinating journalist on the main screen, and the map and source journalists in the thumbnails. From there onwards
[0030] DETAILED DESCRIPTION OF INVENTION
[0031] This invention is a structure of an information retrieval system and a methodology including the constructing a 3D space from multiple images or videos taken of that space from multiple perspectives. Thereafter, introducing icons applications in that constructed 3D space representing access to information related to those points of the icons in that 3D space. Such related information can be anything related to that particular point in space.
[0032] However, one use case of special interest is where the information accessed at that location of the icon by selection, is a video taken from that physical location. In addition, in some embodiments, the icons could have orientation information for the direction of orientation of the cameras taking such videos.
[0033] In some embodiments such videos that are accessed may be related to the images and videos used to construct the 3D space.
[0034] There can be in some embodiments multiple metadata elements, and parametric information also accessed related to such accessed videos. A special case of the use of the invention is in the use in the news industry. Here, consumers of news will in some embodiments have access to the 3D rendering in some embodiments accompanying news stories, so that the news consumer can see the space that is the target of the news story. He or she can navigate around that 3-D space, and understand the context from such navigation. The icons provided at different points on the 3D space when selected will give in some embodiments access for the news consumer, the videos created in those positions of the icons. Therefore, this results in the ability of the user / consumer of news, to "drill down" to videos that are the basis for the news story ("On the spot" feature) . For example, the journalist or newsroom would construct the story from multiple videos of an event, and allow the user / news consumer to access with their own agency any of the source videos that were used for constructing the story, and in some embodiment, construction of the 3D space. In some such embodiments multiple concurrent video streams may be played in secondary windows to aid event verification. This is a new dimension to the news consumption experience where users do not depend on media entities to decide on what they see and don't see, but rather gives them the agency to choose what sources they want to explore and understand as a basis for the news story or breaking story.
[0035] Such a "drill down" feature in some embodiments in news media could be a higher level of subscription service.
[0036] Yet another embodiment, provides features for not just the "drill down", but also access to one-on-one direct dialog with sources creating such video resources for one-on-one communication for in some embodiments, an additional fee.
[0037] In yet another embodiment of this invention, rather than have a 3D space that allows agency for the user to navigate to choose the icons as specified above, instead a map is provided using the geo locations of each of the sources, for the local region around the multiple icons that may represent in some cases video sources that may be selected with the icons for viewing by the user.
[0038] In yet another embodiment rather than have a drill down to the information related to an icon such as a video from that particular location, the user is presented with a map of the local region along with the video side-by-side, so that having visual access to both the map and the video, the user is able to visualize the spacing the local region that is depicted by the map. Such maps in some embodiments are dynamic in that as the location of the video source moves in the field, the icon moves with it. For example the video source moves 5 feet north, the location of the icon will also move by 5 feet to the North. This will give the user a clear understanding of the local environment and what is captured in the videos. In particular, an embodiment for the use case for the news industry, will be extremely powerful where not just the video, but also the location from which the video is taken currently selected icon for this clip of video (or image), so that there is cognitive reinforcement of context of the video information (or images) and the user has a better understanding of the context of the videos that are represented. Yet other embodiments do not give the user the agency to drill down to the specific videos or other information that he or she chooses by selecting the icon. Rather, a set story is constructed with the multiple videos to create a video story, and clips of each of the multiple videos are shown along with highlighting of the icon that is representing the current location of that video. This is depicted in figure 27-11. Still other embodiments have thumbnails of live video and / or the map, (Fig 27-16, 27-17 ) where any of the map, the views of the Sources on the ground or the new generator / anchor are in thumb nails and one of those thumb nails can be displayed on the main screen at which time the corresponding thumb nail can be removed. Some such embodiments may also have live maps where movement of the Sources are depicted on the map. The concurrent display of the multiple views of the sources will also aid in the verification of the event. Such embodiments will allow switching of any of the thumb nails to be on the main screen. Some such embodiments will also have the Source icon on the map to be selectable and thereby present the selected view on the main screen with the map moved to a thumb nail.
[0039] Such embodiments can be played by a news or information consumer with all the thumbnails and main screen synchronous in time.
[0040] In some embodiments, generation of the video stories are as follows. The user can be a user / coordinator / newscaster / journalist who has the ability to sequentially interview and view the perspectives of each of the sources in turn and construct a story that may be published. Some embodiments enable the coordinator to see the thumbnails of each of the sources, and in some of those embodiments a map of the space inhabited by the sources. By tapping the thumbnails, he can have one-on- one interactions on a two-way voice channel for interviews.
[0041] On conclusion of the news story, in some embodiments he is enabled to publish those stories on social media or on news portals.
[0042] Unique feature of this invention is where the viewer for consumer of the news item that is posted on social media or on news portals has the ability to "drill down" to the source videos as previously described. Fig27- 16, 17,18 illustrates some of the views that the news consumer can use.
[0043] In yet another embodiment, videos may be created by individuals that simply see something useful in their location and wish to capture it on video. Therefore, they record video on their phones. They may not be aware of any relationship of this video with other videos that are in their neighborhood. The network in this invention, allows upload of such videos either as streams in real-time or as recordings on the phone which are then uploaded. Then the network analyzes the Geo locations of each of the videos and arranges them into clusters around different Geo locations. Such clusters may be overlapping. Each cluster will have a "key object" which appears in all the members of the cluster. In some embodiments even without the geolocation information characteristics of the visual representation of the object can demonstrate with a acceptable level of certainty that it is the same object and the clusters generated. As an additional embodiment, the characteristics of such key objects may be pre-defined to be a member of a class of objects and in some embodiments an Al engine used to identify members of the class. That class in turn may be determined in some embodiments, by an Al engine that looks at consumer preferences for events / objects of interest, le further refinement of this embodiment, would identify the most desired subjects for a potential consumer market for such videos. Such a mechanism will be done with an Al engine that has been trained with data of objects and issues that are of interest to the viewing consumer.
[0044] The metadata of the video clips may include geolocation and may include orientation data.
[0045] The network in this invention uses image recognition (well known in the background art) along with an Al engine trained on multiple perspective of images and is therefore enabled to identify different views of the object and define the orientation and in some cases the focal length of the camera of each of the videos. This allows the network to provide to a user or consumer to navigate among multiple perspectives of an event or an object that is common in the field of view of each of the clusters. The network can also use the multiperspective videos (and images if uploaded) to reconstruct the 3D remote space.
[0046] In some embodiments the network app on the phone or other device of the Source user, has a permission that allows the usage of any recording made by the source user to be also channeled to the Network for streaming. Some such embodiments can have a switch available to the user that blocks specific videos from being used by the network, to allow exceptions for not recording.
[0047] The videos now integrated into clusters will be stored on the archives for future reference for the verification of videos that are later presented for authentication by the network.
[0048] Yet another embodiment has the membership of video journalists that can use the map interface of the network, where each available source in represented by an icon on a map of the world that can be zoomed in to any particular region. So they can put together video stories as value added and either post back to the network or sell independently to news portals or consumers. So these stories are packaged and available to social media and news portals for a price.
[0049] There can also be the "on the spot" feature for the packaged video to allow drill down to the source videos.
[0050] Yet another embodiment of the invention is a method for news capture. News is random and so the location and time are not known a priori. A challenge therefore is to have the source member (person or in an alternative embodiment machine responsible for the Source) ready to capture the important news event at the location and time when and where it happens. The instant invention offers a method where the Network's mobile ( or other interface) app is installed into the mobiles or other interfaces of a population in a region, and when anyone sees a useful event that can be of interest and can in some embodiments be monetized (the network will in some enbodiments pay the source for their material ).
[0051] The event of that first source member switching on the camera triggers alerts to those Source members in the immediate neighborhood of the first source member and give them the opportunity to switch on their cameras so that multiple views of possible of the event, and recorded on the network. They can also be in some embodiments an alert to a viewer or a interviewer who gets on a map interface all the icons of live sources, that he can navigate to and have dialogue to understand the event.
[0052] In yet other embodiments, the first source member action to start the camera, alerts the interviewer, who is able to then alert other source members in the neighborhood of the first source member, and thereby create multiple views of the event which he or she will be able to navigate to and conduct interviews sequentially.
[0053] The Figures show this operation in some embodiments.
[0054] Fig 27-12, represents a map with multiple source members with their sources randomly distributed around the space. This is essentially a population that is endowed with the network interface which can be a mobile phone. There is an optional viewer / user that has access to this map.
[0055] Fig 27-13, represents a map with the same multiple source members with their sources as in the previous figure, where one of the source members have seen an event of interest ("first source"), and has switched on his or her camera. (In an alternative embodiment there may be other machine based source user such as a drone or a surveillance camera, that can go to the same processes)
[0056] In an embodiment of the invention, as shown in figure 27-14, a predetermined number of sources with the closest Geo locations to the First Source ("second sources") are notified to switch on their cameras. (Such notification may be accompanied in some embodiments by an automatic switch on at their cameras.)
[0057] In some such embodiments, the second sources see the video from the camera of the first source, thereby understanding the event that needs to be covered. They will also see on their screens in some embodiments a local map with the first source highlighted and their own position highlighted among the other sources that have their cameras on. This enables them to understand how to get to the event of interest.
[0058] Fig 27-15, shows the interaction with the viewer / user, where the viewer / user can have two way communication with any one or more of the first source member and the second source members, while benefiting from the recording of these source streams.
[0059] In alternative embodiments, upon the switching on of the first source camera, the viewer / user is immediately notified with an icon on his map. This allows him to the need for exploring what is going on at that location, and therefore he has the capacity to identify a region around that first source, and notify all the sources that are in that region. Such notification, will alert those sources to switch on their cameras. In some such embodiments, such notification to those second sources will also enable their screens to show one both of: the video of the first source, along with a local map representing the position of the first source, and their own position.
[0060] Fig 27-16, 17 are of a presentation to a News consumer with a dynamic map of sources and the views of each source and a view of the News anchor or News room or journalist collating source information to build a coherent story. Each of these elements of visual data are represented on dynamic thumb nails (static for images). One of the views or the map are depicted on the main screen for careful examination. Selection of any of the thumb nails will put that thumb nail on the main screen and revert the current screen image to a thumb nail. It is also possible in some embodiments to have all the views permanently on the thumb nails and one of them by selection displayed on the main screen. Here Fig 27-16 depicts a dynamic map where the sources can move are tracked with the icons. Selection of the icons put the selected source view on the main screen. Alternatively, the selection of a thumb-nail puts the thumb nail video on the main screen (Image if static)
[0061] During the process of construction of the story by the News Anchor / Journalist, he / sh has access to each of the sources and can have one on one conversations with each one of them by selecting them on the map or using the thumb nail. In some embodiments, the other Sources are muted to avoid excessive noise. The News Anchor / Journalist has voice channels to all the sources concurrently otherwise to coordinate the project.
[0062] CONCLUSIONS, RAMIFICATIONS & SCOPE
[0063] It will become apparent that the present invention presented, provides a new paradigm for implementing virtual navigation and visualization of remote spaces.
Claims
CLAIMS1. Method for capturing context of a remote physical space by:- generating a digital twin of the remote space comprising visual records from two or more locations in that remote space-reconstructing a rendition of the remote space;- placing icons at the locations at which the visual records were generated;-linking each of said icons to the visual records upon selection of any one of said icons;Thereby enabling a user to view said visual records in the context of the location in the reconstructed space.
2. A method as in claim 1, wherein said reconstructed rendition of the remote space is a 3D reconstruction using two or more image frames at the periphery of the remote space.
3. A method as in claim 1, wherein said reconstructed rendition of the remote space is a map of the remote space.
4. A method as in any of the preceding claims wherein the visual record is video.
5. A method as in any of the preceding claims wherein the visual record is an image.
6. A method as in any of the preceding claims wherein the visual record is a hologram.
7. A method of capturing a special context of a remote physical space with a plurality of Source Members distributed therein, wherein each of said sources are enabled with cameras comprising:-identifying a member of a predefined event class, by a first Source Member;-switching on of a camera of corresponding source of said First User;- identifying a set of second Source Members among the Source Members, using pre-determined rules related to the physical distance of each of said second Source Members from the First Source Member-enabling the switching on of cameras of said second Sources.
8. A method as in claim 7, wherein said source is configured for hand held operation.
9. A method of claim 7, wherein said source is configured to be supported by a Machine.
10. A method as in claims 7, 8 or 9 wherein the pre-determined rules comprise a circle of a pre-defined radius.
11. A method as in claims 7, 8 or 9 wherein the pre-determined rules comprise a predefined number of Second Source Members at the closes physical distance from the First Source Member.
12. A method as in claims 7, 8, 9, 10 or 11, wherein locations of said First Source Member and Second Source members are represented by icons on a map of the physical space, and a User, is enabled view the map and representations of the views from one or more of said views of the cameras of the First Source Member and the Second Source Members.
13. A method as in claim 12, wherein said User can sequentially view and communicate with one of the Source Members enabled to have a voice communication interface.
14. A method as in claim 12, wherein said User can view thumb nails of a plurality of views of said Source Members.
15. A method as in claim 14, further enabling said User to view a full-screen rendition of the view of one of the Source Members and further enabling said User to switch the full-screen rendering to other Source Members.
16. A method as in any of claims 7 to 15, wherein said Second Source Members see on their screens, one or both of: the video from the camera of the First Source Member; and a local map with their own position and that of the First Source Member represented by highlighted icons, thereby understanding one or both of: the event that is to be covered; and getting to that event.
17. A method as in any of claims 7 to 16, wherein said User is enabled with an interface for two way communications with any one or more of the First Source Member and the Second Source Members.
18. A method for generating a story by a user, wherein the user is enabled to interview using the 2 way audio communication channels with the Sources and wherein the video generated by the Sources are recorded, and wherein the story comprises the perspective provided by the user along with the video records of each of the Sources for substantially the entire time of the story.
19. A method for experiencing a news story from a news portal wherein the news consumer can view the story and concurrently follow the Source videos and is enabled to drill down to any of the source videos for display on the full screen.
20. A method of constructing a representation of a remote space, by integrating separately sources videos, and using one or more of the geo-locations the visual geometries analyzed by methods of object recognition to create clusters of such videos to represent events in local neighborhoods.
21. A method as in Claim 20 further enabled to detect fake events by using the transformations of multiple pixels in the plurality of images at each time frame across the video streams previously captured to construct a possible range of transformations of each of the examined pixels that are geometrically consistent and thereafter examine the transformations required of the potentially fake image or videos to assess authenticity.
Citation Information
Patent Citations
System and method for obtaining and sharing content associated with geographic information
US20080307311A1
Telelocation: location sharing for users in augmented and virtual reality environments
US20180033208A1
System and method for creating a navigable, three-dimensional virtual reality environment having ultra-wide field of view
US20210329222A1
Systems and methods for generating 3D models from drone imaging
US20220398806A1