Data set genome mapping and comparison

The data set genome mapping and comparison system optimizes data processing between client and server sides, caching visualizations to address processing limitations, thereby enhancing user experience and efficiency in visualizing complex data structures like patent genomes.

WO2026006508A1PCT designated stage Publication Date: 2026-01-02COMPANYGENOMICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/035336
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-26
Filing Date
2025-06-26
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing data visualization systems face challenges in efficiently processing large data sets due to limited processing capabilities of client devices, leading to prolonged generation times and suboptimal user experiences, especially when visualizing complex data structures like patent genomes.

Method used

A data set genome mapping and comparison system that dynamically allocates data processing between client-side and server-side, caching data visualizations to improve processing efficiency and reduce time, allowing robust visualization on devices with limited capabilities.

Benefits of technology

Enhances user experience by reducing processing time and resource consumption, enabling efficient visualization of large data sets on devices like smartphones and tablets, while maintaining data integrity and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025035336_02012026_PF_FP_ABST
    Figure US2025035336_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for receiving, storing, generating, mapping, and comparing data sets representing entity genomes are described and illustrated, as are systems and methods for generating visual representations of such genomes. The data sets can be based on a classification-based data structures. Also, the systems and methods can include retrieval of metadata corresponding to earlier-generated graphical representations of genomes, user-selection of genome portions for further comparison with genomes of other entities, generation of graphical representations of multiple genomes, conversion of data frames containing genomic data into different formats, and generation of distances between genomes of different parties representing degrees of similarity between the genomes of the parties based upon one or more distance metrics.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.013350-0004-WO01 DATA SET GENOME MAPPING AND COMPARISON CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 664,697, filed on June 26, 2024, the entire contents of which is hereby incorporated by reference. FIELD

[0002] The present disclosure relates to data set genome comparison on both the client side and the server side of a data set genome mapping and comparison system. SUMMARY

[0003] Information is critical to knowledge, understanding, and decision making. However, when making decisions, decision makers are often presented with too little information, such that a decision is made based on an incomplete view of an issue, or with too much information, such that the amount of information overwhelms the issue and prevents evaluation of the issue in a meaningful way. Improvements in database technologies allow for collection and storage of large amounts of data. Therefore, the obstacle inhibiting good decision making is often the presence of too much information, not too little. Additionally, data stored in raw format in the databases is usually unintelligible to an individual. The management and organization of data or information and being able to effectively convey or present that data or information to an individual is important to evaluating and understanding an issue.

[0004] Data or information is rarely entirely unstructured. In fact, information and data can often be broken down into a hierarchical structure based on attributes or parameters. This structured or semi-structured data can be compiled into data sets and, based on the structure of the data, represented using one or more data visualizations. A data visualization presents data in easy-to-communicate and easy-to-understand graphic or visual representations. For example, data visualization can take the form of histograms, maps (e.g., radial maps), charts, animations, and the like. A master or parent data set containing all of the data for a given visualization can be dynamically modified based on, for example, user modifications to the visualization received through a user interface providing the data visualization, data added to the visualization, multi- visualization data set compilation, etc. By dynamically modifying data sets and allocating dataAttorney Docket No.013350-0004-WO01 processing of certain tasks to either the client-side or the server-side of a data set genome mapping and comparison system, the operation of both the server-side and the client-side can be improved, as well as the effectiveness of the visualized data sets. For example, by pushing some data processing from the server-side to the client-side, the server is capable of handling additional traffic or running more computationally expensive programs based on data received back from the client-side. Additionally, by dynamically modifying data sets on the client-side, the client device limits the amount of data (e.g., as a subset of the master data set) that the client device needs to process to generate a visualization based on the data set. As a result, a robust dynamic visualization of data sets is achieved from a client-side device with relatively limited processing capability (e.g., a smartphone, a tablet, etc.).

[0005] The processing time for generating a data visualization depends on the amount of data being processed. Generating data visualizations on the fly for large amounts of data may require significant processing capability and time. In many cases, the same data may be visualized multiple times in different active sessions of a data set genome mapping and comparison system. In these situations, a data visualization generated during a first active session may be cached such that the data visualization can be reproduced during subsequent active sessions, thereby saving processing resources and time. This significantly improves the user experience when interacting with a data set genome mapping and comparison system.

[0006] In some embodiments, a method is provided that comprises generating, using an electronic processor, on a user interface, a first graphical representation of a number of patent documents assigned to an entity during a first active session, the first graphical representation illustrating the number of patent documents assigned to the entity in a plurality of patent classifications; storing, using the electronic processor, metadata corresponding to the first graphical representation in relation with the entity; receiving, using the electronic processor, a selection of the entity during a second active session; retrieving, using the electronic processor, the metadata stored during the first active session in response to receiving the selection of the entity during the second active session; and generating, using the electronic processor, on the user interface, a second graphical representation based on the metadata.

[0007] Some embodiments disclosed herein provide a method comprising: generating, using an electronic processor, on a user interface, a graphical representation of a number of patentAttorney Docket No.013350-0004-WO01 documents assigned to an entity, the graphical representation illustrating a number of patent documents for a plurality of patent classifications for the entity; receiving, using the electronic processor via the user interface, a selection of a section of the graphical representation, the section covering a subset of patent classifications of the plurality of patent classifications; determining, using the electronic processor, a patent profile of the entity in response to receiving the selection of the section of the graphical representation, the patent profile generated based on a number of patent documents assigned to the entity for the subset of patent classifications; and generating, using the electronic processor, on the user interface, a list of entities ranked based on similarity of patent profile with respect to the patent profile of the entity.

[0008] In some embodiments, a method is provided that comprises receiving, using an electronic processor via a user interface, a selection of a first entity; determining, using the electronic processor, a patent profile of the first entity based on a number of patent documents assigned to the first entity for a plurality of patent classifications; generating, using the electronic processor, on the user interface, a list of entities ranked based on a similarity of a patent profile with respect to the patent profile of the first entity; receiving, using the electronic processor, via the user interface, a selection action of a second entity from the list of entities, the selection action received with respect to the first entity; and generating, using the electronic processor, a graphical representation of a number of patent documents assigned to the first entity and a number of patent documents assigned to the second entity in response to receiving the selection action, the graphical representation illustrating the number of patent documents assigned to the first entity and the number of patent documents assigned to the second entity for the plurality of patent classifications.

[0009] Some embodiments disclosed herein provide a method comprising: querying, using an electronic processor, a patent database based on a time domain and one or more patent classifications; receiving, using the electronic processor, a long-form data frame in response to querying the patent database; converting, using the electronic processor, the long-form data frame into a wide-form data frame by pivoting the long-form data frame; and storing, using the electronic processor, the wide-form data frame in a local memory.

[0010] In some embodiments, a method is provided that comprises querying, using an electronic processor, a patent database; receiving, using the electronic processor, a data frame inAttorney Docket No.013350-0004-WO01 response to the query, the data frame including a patent profile for each of a plurality of patent parties, wherein the patent profile for each of the plurality of patent parties includes a number of patent documents associated with the each of the plurality of patent parties in each of a plurality of patent classifications; receiving, using the electronic processor, via a user interface, a selection of a patent party from the plurality of patent parties; in response to receiving the selection generating, using the electronic processor, on the user interface, a graphical representation illustrating the patent profile of the patent party from the data frame; determining, using the electronic processor, a plurality of distances based on a distance metric, each of the plurality of distances being a distance between the patent party and each other one of the plurality of patent parties represented in the data frame; and generating, using the electronic processor, on the user interface, a ranked list of the plurality of patent parties based on the plurality of distances..

[0011] Before embodiments are explained in detail, it is to be understood that the application of the embodiments disclosed herein is not limited to the details of the configuration and arrangement of components set forth in the following description or illustrated in the accompanying drawings. The techniques described herein are capable of other embodiments and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description, and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof herein are meant to encompass the items listed thereafter and equivalents thereof, as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly, and encompass both direct and indirect mountings, connections, supports, and couplings.

[0012] In addition, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects of the invention may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more processing units, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software based devices, as well as a plurality of different structural components,Attorney Docket No.013350-0004-WO01 may be utilized to implement the invention. For example, “servers” and “computing devices” described in the specification can include one or more processing units, one or more computer- readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the components.

[0013] Other aspects will become apparent by consideration of the detailed description and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] FIG.1 illustrates a data system according to an example embodiment.

[0015] FIG.2 illustrates a server-side processing device according to an example embodiment.

[0016] FIG.3 illustrates a client-side processing device according to an example embodiment.

[0017] FIG.4 illustrates an example client-side interface according to an example embodiment.

[0018] FIG.5 illustrates an example client-side interface according to an example embodiment.

[0019] FIG.6 illustrates an example client-side interface according to an example embodiment.

[0020] FIG.7 illustrates an example client-side interface according to an example embodiment.

[0021] FIG.8 illustrates an example client-side interface according to an example embodiment.

[0022] FIG.9 illustrates an example client-side interface according to an example embodiment.

[0023] FIG.10 illustrates an example client-side interface according to an example embodiment.Attorney Docket No.013350-0004-WO01

[0024] FIG.11 illustrates an example client-side interface according to an example embodiment.

[0025] FIG.12 illustrates an example client-side interface according to an example embodiment.

[0026] FIG.13 illustrates a flowchart of a method for data set genome mapping and comparison according to an embodiment.

[0027] FIG.14 illustrates a flowchart of a method for data set genome mapping and comparison according to an embodiment.

[0028] FIG.15 illustrates a flowchart of a method for data set genome mapping and comparison according to an embodiment.

[0029] FIG.16 illustrates a flowchart of a method for data set genome mapping and comparison according to an embodiment.

[0030] FIG.17 illustrates a flowchart of a method for data set genome mapping and comparison according to an embodiment.

[0031] FIG.18 illustrates a process flow of a method for data set genome mapping and comparison according to an embodiment.

[0032] FIG.19 illustrates an example long-form data frame according to an embodiment.

[0033] FIG.20 illustrates an example wide-form data frame according to an embodiment.

[0034] FIG.21 illustrates an example data matrix according to an embodiment.

[0035] FIG.22 illustrates a process flow of a method for data set genome mapping and comparison according to an embodiment.

[0036] FIG.23 illustrates a process flow of a method for data set genome mapping and comparison according to an embodiment.

[0037] FIG.24 illustrates a flowchart of a method for data set genome mapping and comparison according to an embodiment. DETAILED DESCRIPTIONAttorney Docket No.013350-0004-WO01

[0038] A genome generally refers a complete set of genes for an organism. As used herein, a patent genome similarly refers to a complete set of patent documents for an entity. Patent documents, as used herein, can include published patent documents (whether pending or abandoned), issued patents (where alive, lapsed, or expired), and may cover design patents, utility patents, utility models, invention patents, plant patents, and the like available in patent databases (e.g., USPTO, EPO, etc.). The patent genome for an entity may be limited by a jurisdiction, for example, US (i.e., USPTO), Europe (i.e., EPO), China (i.e., CNIPA), or the like. In the examples used herein, the patent genome may be limited to US, Europe, or both. The patent genome for an entity may also be limited by dates, and / or status, for example, all patent documents filed after a certain date, within a date range, patent documents that are active, or the like. In the example used herein, the patent genome may not be limited by dates and / or status. A data set genome refers to the complete set of available data for a particular category that can be classified.

[0039] Patent documents include inventors who conceived the invention disclosed in the patent document. The patent document may also be assigned to an organization (e.g., company, person, university, government organization, and the like, all referred to herein as an entity). Patent documents also include an Applicant, usually an organization but could be an individual, that filed the patent document. In many cases, the Applicant is also the assignee of the patent document. An entity may refer to the organization to which the patent document is assigned. If the patent document is not assigned, the entity may refer to a set of inventors defined by one or more inventors, which can be referred to as an inventive entity. The patent document is part of the patent genome of the organization to which the patent document is assigned or of the inventive entity of the patent document.

[0040] Patent documents at each jurisdiction’s Patent Office are categorized using classification systems, for example, cooperative patent classification (CPC), international patent classification (IPC), United States patent classification (USPC), and / or the like. The classification systems assign codes to each patent document based on the technology area of the patent document. The classification system code is usually hierarchical. For example, the CPC code includes: (i) section; (ii) class; (iii) subclass; (iv) group; and (v) main group in order of hierarchy. In the following description, examples are described with respect to the CPC system, however, the examples are just as applicable to other classification systems.Attorney Docket No.013350-0004-WO01

[0041] FIG.1 illustrates a data set genome mapping and comparison system 100 operable or configured to receive, store, generate, map, and compare data sets to generate a visual representation (i.e., visualization) of the data set based on a classification-based data structure of the data set. The data set genome mapping and comparison system 100 includes a plurality of client-side devices 105-125, a communication network 130, a server-side mainframe computer or server 145, a first database 150, and a second database 155. The plurality of client-side data input devices 105-125 include, for example, a server 105, a personal computer 110, a tablet computer 115, a personal digital assistant (“PDA”) (e.g., an e-reader, etc.) 120, and a mobile phone (e.g., a smart phone) 125. The above noted client-side devices 105-125 are provided as an example. Other types of devices may be used as client-side devices in the data set genome mapping and comparison system 100. Each of the devices 105-125 is operable or configured to communicatively connect to the server 145 through the communication network 130 and receive a data set for comparison from the server 145. The data sets can be received from the databases 150, 155 for visualization. All of the components of the data set genome mapping and comparison system 100 may not be maintained by the same entity, and may be distributed between different entities in various configurations.

[0042] The communication network 130 may include, for example, a wide area network (“WAN”) (e.g., a TCP / IP based network), a local area network (“LAN”), a neighborhood area network (“NAN”), a home area network (“HAN”), or personal area network (“PAN”) employing any of a variety of communications protocols, such as Wi-Fi, Bluetooth, ZigBee, etc. The communication network 130 may also or alternatively include is a cellular network, such as, for example, a Global System for Mobile Communications (“GSM”) network, a General Packet Radio Service (“GPRS”) network, a Code Division Multiple Access (“CDMA”) network, an Evolution-Data Optimized (“EV-DO”) network, an Enhanced Data Rates for GSM Evolution (“EDGE”) network, a 3GSM network, a 4GSM network, a 4G LTE network, a Digital Enhanced Cordless Telecommunications (“DECT”) network, a Digital AMPS (“IS-136 / TDMA”) network, or an Integrated Digital Enhanced Network (“iDEN”) network, etc.

[0043] The connections between the devices 105-125 and the communication network 130 are, for example, wired connections, wireless connections, or a combination of wireless and wired connections. Similarly, the connections between the server 145 and the communication network 130 are wired connections, wireless connections, or a combination of wireless and wiredAttorney Docket No.013350-0004-WO01 connections. In some embodiments, the communication network 130, and / or the communications between the devices 105-125 and the server 145, are protected using one or more encryption techniques, such as those techniques provided in the IEEE 802.1 standard for port-based network security, pre-shared key, Extensible Authentication Protocol (“EAP”), Wired Equivalency Privacy (“WEP”), Temporal Key Integrity Protocol (“TKIP”), Wi-Fi Protected Access (“WPA”), etc.

[0044] FIG.2 illustrates the server-side of the data set genome mapping and comparison system 100 with respect to the server 145. The server 145 is electrically and / or communicatively connected to a variety of modules or components of the data set genome mapping and comparison system 100. For example, the illustrated server 145 is connected to the first database 150, and the second database 155. The server 145 includes a controller 200, a power supply module 205, and a network communications module 210. The controller 200 includes combinations of hardware and software that are operable to, among other things, dynamically modify data sets on the server side of the data set genome mapping and comparison system 100. In some constructions, the controller 200 includes a plurality of electrical and electronic components that provide power, operational control, and protection to the components and modules within the controller 200 and / or the data set genome mapping and comparison system 100. For example, the controller 200 includes, among other things, a processing unit 215 (e.g., a microprocessor, a microcontroller, or another suitable programmable device), a memory 220, input units 225, and output units 230. Also, in some constructions, the processing unit 215 includes, among other things, a control unit 235, an arithmetic logic unit (“ALU”) 240, and a plurality of registers 245 (shown is a group of registers in FIG.2), and is implemented using a known computer architecture, such as a modified Harvard architecture, a von Neumann architecture, etc. The processing unit 215, the memory 220, the input units 225, and the output units 230, as well as the various modules connected to the controller 200 are connected by one or more control and / or data buses (e.g., common bus 250). The control and / or data buses are shown schematically in FIG.2 for illustrative purposes. The use of one or more control and / or data buses for the interconnection between and communication among the various modules and components would be known to a person skilled in the art in view of the invention described herein.Attorney Docket No.013350-0004-WO01

[0045] The memory 220 includes, for example, a program storage area and a data storage area. The program storage area and the data storage area can include combinations of different types of memory, such as read-only memory (“ROM”), random access memory (“RAM”) (e.g., dynamic RAM [“DRAM”], synchronous DRAM [“SDRAM”], etc.), electrically erasable programmable read-only memory (“EEPROM”), flash memory, a hard disk, an SD card, or other suitable magnetic, optical, physical, electronic memory devices, or other data structures. The processing unit 215 is connected to the memory 220 and executes software instructions that are capable of being stored in a RAM of the memory 220 (e.g., during execution), a ROM of the memory 220 (e.g., on a generally permanent basis), or another non-transitory computer readable data storage medium such as another memory or a disc.

[0046] In some embodiments, the controller 200 or network communications module 210 includes one or more communications ports (e.g., Ethernet, serial advanced technology attachment [“SATA”], universal serial bus [“USB”], integrated drive electronics [“IDE”], etc.) for transferring, receiving, or storing data associated with the data set genome mapping and comparison system 100 or the operation of the data set genome mapping and comparison system 100. Software included in the implementation of the data set genome mapping and comparison system 100 can be stored in the memory 220 of the controller 200. The software can include, for example, firmware, one or more applications, program data, filters, rules, one or more program modules, and other executable instructions. The controller 200 is configured to retrieve from memory and execute, among other things, instructions related to the dynamic data modification processes and methods described herein. In other constructions, the controller 200 includes additional, fewer, or different components.

[0047] The power supply module 205 supplies a nominal AC or DC voltage to the controller 200 or other components or modules of the data set genome mapping and comparison system 100. The power supply module 205 is powered by, for example, mains power having nominal line voltages between 100V and 240V AC and frequencies of approximately 50-60Hz. The power supply module 205 can also be operable or configured to supply lower voltages to operate circuits and components within the controller 200 or the data set genome mapping and comparison system 100. In other constructions, the controller 200 or other components and modules within the data set genome mapping and comparison system 100 are powered by one orAttorney Docket No.013350-0004-WO01 more batteries or battery packs, or another grid-independent power source (e.g., a generator, a solar panel, etc.).

[0048] FIG.3 illustrates the client-side of the data set genome mapping and comparison system 100 with respect to the client-side devices 105-125. The client-side devices 105-125 include a controller 300, a power supply module 305, a network communications module 310, a user interface 315, and a client-side database 320. The controller 300 includes combinations of hardware and software that are configured to, among other things, dynamically modify data sets on the client side of the data set genome mapping and comparison system 100. In some constructions, the controller 300 includes a plurality of electrical and electronic components that provide power, operational control, and protection to the components and modules within the controller 300 and / or the data set genome mapping and comparison system 100. For example, the controller 300 includes, among other things, a processing unit 325 (e.g., a microprocessor, a microcontroller, or another suitable programmable device), a memory 330, input units 335, and output units 340. Also in some constructions, the processing unit 325 includes, among other things, a control unit 345, an arithmetic logic unit (“ALU”) 350, and a plurality of registers 355 (shown is a group of registers in FIG.3), and is implemented using a known computer architecture, such as a modified Harvard architecture, a von Neumann architecture, etc. The processing unit 325, the memory 330, the input units 335, and the output units 340, as well as the various modules connected to the controller 300 are connected by one or more control and / or data buses (e.g., common bus 360). The control and / or data buses are shown schematically in FIG.3 for illustrative purposes. The use of one or more control and / or data buses for the interconnection between and communication among the various modules and components would be known to a person skilled in the art in view of the invention described herein.

[0049] The memory 330 includes, for example, a program storage area and a data storage area. The program storage area and the data storage area can include combinations of different types of memory, such as read-only memory (“ROM”), random access memory (“RAM”) (e.g., dynamic RAM [“DRAM”], synchronous DRAM [“SDRAM”], etc.), electrically erasable programmable read-only memory (“EEPROM”), flash memory, a hard disk, an SD card, or other suitable magnetic, optical, physical, electronic memory devices, or other data structures. The processing unit 325 is connected to the memory 330 and executes software instructions that are capable of being stored in a RAM of the memory 330 (e.g., during execution), a ROM of theAttorney Docket No.013350-0004-WO01 memory 330 (e.g., on a generally permanent basis), or another non-transitory computer readable data storage medium such as another memory or a disc.

[0050] In some embodiments, the controller 300 or network communications module 310 includes one or more communications ports (e.g., Ethernet, serial advanced technology attachment [“SATA”], universal serial bus [“USB”], integrated drive electronics [“IDE”], etc.) for transferring, receiving, or storing data associated with the data set genome mapping and comparison system 100 or the operation of the data set genome mapping and comparison system 100. Software included in the implementation of the data set genome mapping and comparison system 100 can be stored in the memory 330 of the controller 300. The software includes, for example, firmware, one or more applications, program data, filters, rules, one or more program modules, and other executable instructions. The controller 300 is configured to retrieve from memory and execute, among other things, instructions related to the dynamic data modification processes and methods described herein. In other constructions, the controller 300 includes additional, fewer, or different components.

[0051] The power supply module 305 supplies a nominal AC or DC voltage to the controller 300 or other components or modules of the data set genome mapping and comparison system 100. The power supply module 305 is powered by, for example, mains power having nominal line voltages between 100V and 240V AC and frequencies of approximately 50-60Hz. The power supply module 305 can also be operable or configured to supply lower voltages to operate circuits and components within the controller 300 or the data set genome mapping and comparison system 100. In other constructions, the controller 300 or other components and modules within the data set genome mapping and comparison system 100 are powered by one or more batteries or battery packs, or another grid-independent power source (e.g., a generator, a solar panel, etc.).

[0052] The illustrated user interface 315 includes a combination of digital and analog input or output devices required to achieve a desired level of control and monitoring for the data set genome mapping and comparison system 100. For example, the user interface 315 includes a display (e.g., a primary display, a secondary display, etc.) and input devices such as touchscreen displays, a plurality of knobs, dials, switches, buttons, etc. The display is, for example, a liquid crystal display (“LCD”), a light-emitting diode (“LED”) display, an organic LED (“OLED”)Attorney Docket No.013350-0004-WO01 display, an electroluminescent display (“ELD”), a surface-conduction electron-emitter display (“SED”), a field emission display (“FED”), a thin-film transistor (“TFT”) LCD, or the like.

[0053] An electronic processor as used herein may refer to one or both of the controllers 200, 300 or one or both of the processing units 215, 325. The functionality of the electronic processor as described herein may be performed by any one of or a combination of the controllers 200, 300 and the processing units 215, 325. The electronic processor may be provided on a single device or distributed across multiple devices, and may be provided in a cloud computing form. The electronic processor may be maintained by a single entity or multiple entities across various platforms.

[0054] The server 145 interacts over the communication network 130 with the various client- side devices 105-125 to receive and service requests for generating dynamic data visualizations based on data sets provided by the server 145. In some examples, the server 145 may also interact with the first and the second database 150, 155 over the communication network 130. In one example, the server 145 may store a data set genome server application in the memory 220. The controller 200 may run or execute the data set genome server application to provide data to the client devices 105-125 related to the visualizations. The client devices 105-125 may similarly store a data set genome client application in the memory 330. The controller 300 may run or execute the data set genome client application to interact with the server 145. When initiated, the data set genome client application attempts to establish a connection between the client device 105-125 and the server 145 by sending a connection request to the data set genome server application. Upon receiving a connection request, the data set genome server application may partition disk space (that is, memory space) to establish the connection between the server 145 and the client device 105-125. Each instantiation of the connection between the server 145 (and therefore the data set genome server application) and the client device 105-125 (and therefore the data set genome client application) may be referred to as an active session. In some examples, the data set genome client application is a web application that is accessed using a web browser executed on a client device. In these examples, the active session is a web session.

[0055] The active session is used to exchange parameters and data for visualizations. For example, for a given visualization, the server 145 provides the data to the client devices 105-125 related to that visualization (e.g., objects or segments of the data visualization, attributes of eachAttorney Docket No.013350-0004-WO01 object and asset, etc.). The server 145 provides this master or parent data set to the client device and allocates the modification and generation of new or different data sets based on the master or parent data set to the client device 105-125. Additions or modifications to the master or parent data set (e.g., including the generation of additional data subsets) can be provided from the client devices 105-125 to the server 145 so the server can update the database and data sets corresponding to the dynamic data visualization. In some embodiments, the server 145 compiles multiple dynamic data visualizations from different users into a single, comprehensive data map (e.g., where the sub-maps are modified by different users and a supervisor is permitted access to the combined data maps and data sets).

[0056] FIG.4 illustrates a client-side interface 400 for data set genome mapping and comparison that may be displayed during an active session of the data set genome client application according to some examples. The illustrated client-side interface 400 includes a navigation pane 405, a search pane 410, and a visualization section 415. In the example illustrated, the navigation pane 405 is provided at a top of the client-side interface 400 above the visualization section 415. The search pane 410 is provided on a side of the visualization section 415 below the navigation pane 405. The navigation pane 405 includes a plurality of tabs 420. The tabs 420 are user selectable to navigate among the tabs 420, each of which includes a different visualization section 415. In some examples, the navigation pane 405 and the search pane 410 are frozen or locked such that the navigation pane 405 and the search pane 410 are visible regardless of the tab 420 selected. In some examples, the plurality of tabs includes a patents overview tab 420A, a patents detail tab 420B, a nearest neighbor tab 420C, a find neighbors tab 420D, a compare neighbors tab 420E, and / or a neighbor clusters tab 420F.

[0057] The illustrated search pane 410 includes a search section 425 and a settings section 430. The illustrated search section 425 includes a search dropdown 435. In response to receiving user input selecting the search dropdown 435, a search sub-pane 440 (as shown in FIG.5) is displayed (by the client application) over the visualization section 415. In some embodiments, the search sub-pane 440 includes a search textbox 445 and a search list 450. The search textbox 445 may be used to receive keywords for performing an entity search. Results from the search are displayed in the search list 450. The results in the search list 450 may be sorted based on relevancy or other criteria. Prior to the search, the search list 450 may include a list of previously searched or selected entities, or a list of ranked entities based on number of assigned patents, theAttorney Docket No.013350-0004-WO01 similarity between the searched entity name and other entity names, or other criteria. A selection (e.g., a user click / touch action) can be received from the search list 450 selecting an entity for visualization.

[0058] As shown in FIG.6, the illustrated settings section 430 includes a plurality of dropdown menus including, for example, a neighbor settings section 455 and a global settings section 460 to select various settings for the visualization in the visualization section 415. The neighbor settings section 455 is used to select, for example, a neighbor type (e.g., a type of entity), a CPC level, CPC section(s), and CPC code(s). The global settings section 460 can be used to retrieve results of a search based upon a minimum amount of data (e.g., only those entities having a minimum number of patents and / or published patent applications). The global settings section 460 can also or instead be used to control the manner in which search results are generated (e.g., a ranking category by which a patent genome of an entity is compared to those of other entities). The settings section 430 may include additional or fewer options for selection.

[0059] FIG.7 illustrates a visualization section 465 when the patents overview tab 420A is selected. The visualization section 465 includes a plurality of visualization subsections including, for example, a patents over time subsection 470, a top CPC codes subsection 475, a patents by party type subsection 480, and a patents by CPC codes subsection 485. The patents over time subsection 470 may display visualizations or listings relating to patent filings over time, for example, over years for a selected entity. The patents over time subsection 470 may include two tabs: (i) a CPC section tab; and (ii) a patent types tab. When the CPC section tab is selected, a visualization or listing is displayed showing the number of patent documents by year and CPC section. When the patent types tab is selected, a visualization or listing is displayed showing the number of patent documents by year and patent types (e.g., utility, design, reissue, utility model, and the like).

[0060] The top CPC codes subsection 475 may display visualizations or listings relating to top CPC codes for patent filings and / or patent grants of the entity. The top CPC codes subsection 475 may include multiple tabs that display the patent filings and / or patent grants by CPC section, CPC subsections, CPC groups, and CPC subgroups. The patents party type subsection 480 may display visualizations or listings relating to assignees, inventive entities, or other entities (e.g., law firms) relating to the patent filings and / or patent grants. The patents by CPC codesAttorney Docket No.013350-0004-WO01 subsection 485 may display visualizations or listings related to patent documents by CPC codes for the selected entity. The patents by CPC codes subsection 485 may include multiple tabs that display the patent filings and / or patent grants by CPC section, CPC subsections, CPC groups, and CPC subgroups. The visualization section 465 may include more or fewer subsections. For example, additional subsections may display the name of the selected entity, the total number of patent documents for the selected entity including the date of the first patent document and a date of the latest patent document, a top plurality of CPC codes and the number of patent documents for the entity in the top plurality of CPC codes, or any combination thereof.

[0061] FIG.8 illustrates a visualization section 490 when the patents detail tab 420B is selected. The illustrated visualization section 490 is a sortable listing of the patent genome, that is, all the patent documents of the selected entity. The sortable listing may include columns displaying an assignee, a law firm, a patent document number, a date relating to the patent document number (e.g., patent application filing date, patent application publication date, patent grant date, and the like), one or more hierarchical codes of the CPC (e.g., sections, subsection, classes, etc.), or any combination thereof. In some embodiments, the illustrated visualization section 490 is a subset of the patent genome, such as if the results displayed are limited by a date range using one or more user-selectable filters (not shown).

[0062] FIG.9 illustrates a visualization section 495 when the nearest neighbor tab 420C is selected. The illustrated visualization section 495 includes a Patent Genome subsection 500, a patents by CPC codes subsection 485, and a nearest neighbor subsection 510. The Patent Genome subsection 500 displays a visualization of the patent genome of the selected entity. Specifically, the patent genome subsection 500 illustrates the number of patents by CPC code class showing all of the existing CPC code classes. The nearest neighbor subsection 510 displays a table showing a list of entities corresponding to the selected entity, ranked by the similarity of each entity’s patent genome to that of the selected entity. The list of entities is ranked according to a similarity score as described in more detail below. Methodologies for determining the similarity score for an entity are provided below. In some embodiments, the illustrated visualization section 495 is a subset of the patent genome, such as if the results displayed are limited by a date range using one or more user-selectable filters (not shown), or by particular CPC sections or CPC codes in the neighbor settings section 455 of the settings section 430.Attorney Docket No.013350-0004-WO01

[0063] The illustrated nearest neighbor subsection 510 displays a nearest neighbor listing table providing a ranked list of entities corresponding to the selected entity. The list of entities is ranked according to a similarity score. The nearest neighbor listing table may include columns: (i) rank; (ii) entity name; and (iii) similarity score of some or all of the entities in the table. The rank column and the similarity score column provide the rank and the similarity score respectively for the corresponding entity listed in the entity name column. The nearest neighbor listing table may be sortable by selecting one of the above three columns.

[0064] FIG.10 illustrates a visualization section 515 when the find neighbors tab 420D is selected. The illustrated visualization section 515 includes a Party Patents by CPC Codes subsection 520, a Neighbor Patents by CPC Codes subsection 525, and a Neighbors subsection 510. The Neighbor Patents by CPC Codes subsection 525 may display visualizations or listings related to patent documents by CPC codes section for a neighbor entity. The neighbor entity may be automatically selected as the lowest ranked neighbor in the Neighbor subsection 510 (e.g., the neighbor having a patent genome with the greatest calculated similarity to that of the selected entity illustrated in the Party Patents by CPC Codes subsection 520). If a different neighbor entity is selected, for example, by a user action received through the client application (e.g., selecting a different entity in the nearest neighbor subsection 510), the Neighbor Patents by CPC Codes subsection 525 displays visualizations or listings related to patent documents by CPC codes section for the selected neighbor entity.

[0065] The Patents by CPC Codes subsection 520 and the Neighbor Patents by CPC Codes subsection 525 illustrated in FIG. 10 each include two tabs: (i) a table tab (e.g., a first tab); and (ii) a plot tab (e.g., a second tab). The table tab and the plot tab are user selectable to display the corresponding visualizations. When the table tab is selected, a patent listing table is displayed showing the number of patent documents assigned to the selected entity by CPC code. The patent listing table may include four columns: (i) assignee name; (ii) CPC code; (iii) description of CPC code (or title of CPC code); and (iv) number of patent documents. The assignee name displays the name of the selected entity or a sub-entity (e.g., a subsidiary, holding company, or the like) of the selected entity to which the patents and / or patent applications in the particular CPC code are assigned. The CPC code and the description of CPC code columns may only display the CPC codes which include patent documents assigned to the selected entity. The number of patent documents column displays the number of patent documents assigned to the selected entity forAttorney Docket No.013350-0004-WO01 the particular CPC code. The patent listing table may be sortable by selecting one of the above four columns.

[0066] When the plot tab is selected (as shown in FIG.10), a patent listing visualization illustrating the number of patent documents assigned to the selected entity by CPC code is displayed. The patent listing visualization may be a histogram with CPC codes provided on the x-axis, and the number of patent documents provided on the y-axis (or vice-versa). The histogram bars correspond to the number of patent documents for each of the CPC codes. Specifically, the patent listing visualization provides a visualization of the patent listing table displayed in the table tab. In other examples, the patent listing visualization may include a different visualization, for example, bar charts, pie charts, and the like.

[0067] In the example shown in FIG.10, the table of the Neighbors subsection 510 may display a plurality of similarity scores calculated according to various methods. The Neighbors table shown in FIG.10 may include columns: (i) entity type (e.g., patent Assignee, patent Applicant, etc.); (ii) entity name; (iii) overall scores; and (iv) ranking by metric. The nearest neighbor listing table 510 may be sortable by selecting any of the columns shown.

[0068] In some embodiments, the Patent Genome subsection 500 of FIG 9 and the Party Patents by CPC Codes subsection 520 of FIG.10 illustrate the same data: the number of patent documents assigned to an entity by CPC code class for a range of CPC code classes or for all CPC code classes. In the illustrated embodiment, all CPC code classes are illustrated. In other embodiments, a subset of CPC code classes is illustrated. It can be advantageous to illustrate less than all CPC code classes for several reasons. For example, it may be difficult or impractical to illustrate all CPC code classes in these subsections 500, 520, such as when the there are hundreds or even thousands of available CPC code classes. In the illustrated embodiment, it is possible to illustrate all CPC code classes at the group level (at the bottom of the chart illustrated in these subsections 500, 520). However, when presenting CPC code classes at the sub-group level, there are many more CPC code classes that may not be possible to easily display in a similar manner. As another example, it may be advantageous to display only a particular range or ranges of available CPC code classes, such as only CPC code classes having to do with certain sets or fields of technology (e.g., piston pumps, semiconductor chip fabrication, etc.).Attorney Docket No.013350-0004-WO01

[0069] With continued reference to subsections 500, 520 illustrating the number of patent documents assigned to the entity by CPC code at the subsection level, there are a number of CPC code classes in which the entity has no patent documents assigned to the entity. Therefore, the lack of patent documents in certain CPC code classes can carry just as much valuable information as the existence of patent documents for the entity in other CPC code classes.

[0070] It is often desirable to compare the Patent Genome 500 or the Party Patents by CPC Codes subsection 520 (terms that will hereinafter be used interchangeably) of an entity with that of another entity, or even with those of multiple entities. This comparison can be carried out in a number of different ways. In some methods described herein by way of example only, the Patent Genome 500 is a data set including dimensions defined at least in part by CPC code classes (whether at the section, sub-section, group, sub-group, or other level) and magnitudes defined at least in part by the number of patent documents assigned to the entity in each CPC code class. Comparing such a dimensional data set of one entity with that of another can be performed by any of the distance metrics as further described with respect to FIGs.19-21 that generally determine or quantify a distance between data sets.

[0071] Still other methods for comparing the data set of a Patent Genome 500 having at least two dimensions can also or instead be used. For example, the data set of the Patent Genome 500 illustrated in FIG. 9 can be compared to any number of other data sets of other entities by running the image of the Patent Genome 500 shown in FIG.9 through a number of well-known image recognition engines more commonly used for identifying similarities between photographs or other images, and for ranking a plurality of images based upon their similarity to one another. As another example, various metrics used to characterize graphs (e.g., statistical graphs) may be used to compare Patent Genome data sets, such as, for example, comparing shapes or skews, outliers, centers or medians, spreads, or a combination thereof. Similarly, curve fitting may be used to represent a line graph of a histogram (e.g., connecting magnitudes over the classes) and defined curves (or functions) or components or parameters of the same may be compared to compare the data set of a Patent Genome 500 with one more other data sets of other Patent Genomes. Areas under such curves may also be determined and compared to compare the data set of a Patent Genome 500 with one more other data sets of other Patent Genomes. One of ordinary skill in the art will recognize that these and other methods can be utilized to compareAttorney Docket No.013350-0004-WO01 the Patent Genome data sets as described and illustrated herein. These alternative methods all fall within the spirit and scope of the present invention.

[0072] After the Patent Genome of an entity has been compared to that of one or more other entities, the resulting ranking can be illustrated (e.g., in nearest neighbor subsection 510 of FIG. 10), such as in a listing showing the type, name, and rank of each entity as described above. In the illustrated embodiment by way of example only, the rank of each entity calculated by Euclidean, Manhattan, Chebyshev, and Angular comparison methods is presented in a table as shown. Rather than rely upon any one comparison method, in some embodiments the mean of all ranks for each entity (e.g., Euclidean, Manhattan, Chebyshev, and Angular comparison methods shown in FIG. 10) is provided as shown in the “Mean” column of the “Overall” section in FIG. 10, with the corresponding standard deviation between the calculated ranks also shown in the “SD” column of the “Overall” section in FIG.10. The ranks calculated by any number of comparison methods can be presented in the Neighbors subsection 510.

[0073] In some embodiments, the Neighbors subsection 510 can be reconfigured to display the entities by any of the comparison methods, such as by user selection of any of the “Ranking by Metric” headers. Also in some embodiments, a radio button or other user-manipulatable control may be provided for each row such that the corresponding entity may be selected for display in the Neighbor Patents by CPC Codes subsection 525.

[0074] FIG.11 illustrates a visualization section 530 displayed when the compare neighbors tab 420E (see FIG.4) is selected. The visualization section 530 includes a Party and Neighbor Genomes subsection 535 and a Patents by CPC codes subsection 540. The Party and Neighbor Genomes subsection 535 and the Patents by CPC codes subsection 540 may display visualizations or listings related to the patent genomes classified according to CPC code subsections for the selected entity, and a plurality of neighbors from the ranked list of neighbors for the selected entity (e.g., shown in FIGS.9 and 10). The plurality of neighboring entities may be automatically selected based on predetermined criteria, for example, similarity rank, similarity score, or the like. The visualizations provided in the Party and Neighbor Genomes subsection 535 may be similar to the visualization displayed in the plot tab of the Neighbor Patents by CPC Codes subsection 525 (see FIG.10). The Patents by CPC codes subsection 540 lists the numberAttorney Docket No.013350-0004-WO01 of patents by CPC codes for the plurality of nearest neighbors that are visualized in the Party and Neighbor Genomes subsection 535.

[0075] FIG.12 illustrates a visualization section 545 when the neighbor clusters tab 420F is selected. The visualization section 545 includes a neighbors slider 550, a clusters slider 555, a clustering method selector 560, a Dendrogram subsection 565, and a Neighbors by Cluster subsection 570. The neighbors slider 550 includes a slider bar to select a number of neighbors to be displayed in the Dendrogram subsection 565. The clusters slider 555 includes a slider bar to select a number of clusters to display in the Dendrogram subsection 565. The clustering method selector 560 includes a drop down menu to select a clustering method from a list of clustering methods. An example methodology for generating clusters is described below with respect to FIG.20. The dendrogram subsection 565 displays a dendrogram of the selected entity and the selected number of nearest neighboring entities corresponding to the selected entity. The Neighbors by Cluster subsection 570 displays a table showing nearest neighbors of the selected entity with respect to their clusters.

[0076] It should be noted that the various user interfaces illustrated and described herein with respect to FIGS.4-12 are examples only, and that the various sections of the user interfaces may be arranged and sized in various alternative manners and configurations.

[0077] FIG.13 is a flowchart of an example method 600 for data set genome mapping and comparison. The method 600 is described as being performed by the server 145, and in particular the controller 200 included in the server, for example, when executing the data server genome server application. Accordingly, the electronic processor as referred to with respect to method 600 may refer to the controller 200 and corresponding processing unit 215. In some examples, the method 600 may be distributed between the controller 200 of the server 145 and the controller(s) 300 of one or more of the client devices 105-125, and between the data server genome server application and data server genome client application. In the example illustrated, the method 600 includes receiving, using an electronic processor, via a user interface 315, a selection of a first entity during an active session (at block 610). The active session may be initiated when a client device 105-125 establishes a connection with the server 145. The server 145 may assign a session identifier (ID) for the active session. The active session may include a timed session where the active session is active until the end of a time period, at which time theAttorney Docket No.013350-0004-WO01 session ID expires. The selection of the first entity may be received via the search section 425. A user may search for the entity using the search textbox 445 and select the desired entity from the search list 450 as the first entity.

[0078] The method 600 also includes determining, using the electronic processor, a patent profile of the first entity (at block 620). The patent profile is determined based on a number of patent documents assigned to the first entity for a set of patent classifications. The patent profile is based on the patent genome of the first entity. The electronic processor determines a patent profile by determining the number of patent documents assigned to each classification system code in the patent genome. The electronic processor may determine the number of patent documents assigned to each of the sections, sub-sections, groups, main groups, or other classifications of a classification system. In one example, the patent profile is determined on the fly in response to receiving the selection of the first entity. In another example, the patent profile is pre-determined for each entity, and the patent profile for the first entity is retrieved from the pre-determined profiles in response to receiving the selection of the first entity. As time goes on, each entity may experience a change in patent documents due to new filings, change of assignees, acquisitions, and the like. The patent profiles may be periodically updated based on the availability of new data.

[0079] The illustrated method 600 further includes generating, using the electronic processor, on the user interface 315, a first graphical representation of a number of patent documents assigned to the first entity (at block 630). The first graphical representation illustrates a number of patent documents for each patent classification in the set of patent classifications for the first entity. In one example, the first graphical illustration is a histogram, and includes classes of the classification system code on the X-axis, and a number of patent documents on the Y-axis. The first graphical illustration may include different types of visualizations as discussed above. One example of the first graphical representation is provided by the Patent Genome subsection 500 as shown in FIG.9. The first graphical representation is generated by analyzing the patent profile of the first entity. The first graphical representation may also be displayed in the plot tab of the Party Patents by CPC Codes subsection 520 as shown in FIG.10, and the Party and Neighbor Genomes subsection 535 as shown in FIG.11.Attorney Docket No.013350-0004-WO01

[0080] The method 600 of FIG. 13 further includes generating, using the electronic processor, on the user interface 315, a list of entities ranked based on a similarity of a patent profile of each entity with respect to the first entity (at block 640). The electronic processor compares the patent profile of the first entity with the patent profile of other entities in any of the manners described above with reference to the Patent Genome illustrated in the Patent Genome subsection 500 of FIG.9 and the Party Patents by CPC Codes subsection 520 of FIG.10, and determines the similarity of each entity with the first entity in terms of their patent profiles. Specifically, the electronic processor compares the patent profiles, and generates a similarity score for each comparison as described above. The entity with the higher similarity score is placed higher in the ranked list of entities.

[0081] The method 600 of FIG. 13 further includes generating, using the electronic processor, on the user interface 315, a second graphical representation of a number of patent documents assigned to a second entity from the list of entities (at block 650). The second graphical representation illustrates a number of patent documents for each patent classification for the second entity. In one example, the second graphical illustration is a histogram, and includes classes of the classification system code on the X-axis and a number of patent documents on the Y-axis. The second graphical illustration may include different types of visualizations as discussed above. The second graphical representation is generated by analyzing the patent profile of the second entity. The second graphical representation may be displayed in the plot tab of the Neighbor Patents by CPC Codes subsection 525 as shown in FIG.10, and the Party and Neighbors Genomes subsection 535 as shown in FIG.11. The second entity for generating the second graphical representation may be automatically selected by the electronic processor from the list of entities. For example, the electronic processor may automatically select the top entry in the list of entities. The second entity for generating the second graphical representation may also or instead be selected based on a user selection from the list of entities.

[0082] The above-noted visualizations (that is, graphical representations) are generated based on analyzing the data retrieved from the databases 150 and 155. The amount of time taken to generate a visualization depends on the amount of data to be analyzed. When the visualizations are generated separately for each active session, even if the same visualization was generated in a previous active session for the same client-side device 105-125, the user experience decreases, as the user is having to wait longer amounts of time before the visualization is generated.Attorney Docket No.013350-0004-WO01

[0083] FIG.14 is a flowchart of an example method 700 for data set genome mapping and comparison to improve the user experience. The method 700 is described as being performed by the server 145, and in particular the controller 200 included in the server, for example, when executing the data server genome server application. Accordingly, the electronic processor as referred to with respect to method 700 may refer to the controller 200 and corresponding processing unit 215. In some examples, the method 700 may be distributed between the controller 200 of the server 145 and the controller(s) 300 of one or more of the client devices 105-125, and between the data server genome server application and data server genome client application. In the example illustrated in FIG.14, the method 700 includes generating, using an electronic processor, on a user interface 315, a first graphical representation of the number of patent documents assigned to an entity during a first active session (at block 710). As discussed above with respect to method 600, first and second graphical representations relating to one or more entities may be generated during a first active session. The graphical representations may include the patent genome representations, patents over time representations (e.g., as shown in FIG.7), top CPC codes representations (e.g., as shown in FIG. 7), party or neighbor patents by CPC codes representations (as shown in FIG.10), party and neighbor genome representations (as shown in FIG.11), dendrogram and cluster representations (as shown in FIG.12), or the like. Specifically, the graphical representation includes a number of patent documents assigned to the entity for a plurality of patent classifications.

[0084] The method 700 including storing, using the electronic processor, metadata corresponding to the first graphical representation in relation with the entity (at block 720). The electronic processor stores the metadata relating to the first graphical representation in relation to the entity, for example, in a database. In some examples, the metadata is also stored in relation to user data or a user account. The metadata may be stored in, for example, the memory 330. The electronic processor may use a caching process to save the metadata either in relation to the user account in the memory 220 or in the memory 330. The metadata includes, for example, data points and / or values for quickly and easily regenerating the first graphical representation. For example, when the first graphical representation is a histogram, the metadata may include the number of bars, the height of each bar, and the like. The metadata may store sufficient data such that the first graphical representations can be quickly and easily recreated without the need forAttorney Docket No.013350-0004-WO01 the underlying data that was used to generate the first graphical representation in the first active session.

[0085] The illustrated method 700 also includes receiving, using the electronic processor, vias the user interface 315, the selection of the entity during a second active session (at block 730). The second active session may be initiated when a client device 105-125 establishes a connection with the server 145 that is distinct from the first active session. For example, the second active session may be initiated after the first active session expires. The server 145 may assign a session ID for the second active session that is different from the first active session. The selection of the entity may be received via the search section 425. A user may search for the entity using the search textbox 445 and select the desired entity from the search list 450 as the entity.

[0086] The illustrated method 700 also includes retrieving, using the electronic processor, the metadata stored during the first active session in response to receiving the selection of the entity during the second active session (at block 740). The electronic processor retrieves the metadata stored during the first active session to regenerate the graphical representation. The metadata is retrieved in response to receiving the selection of the entity during the subsequent active session.

[0087] The illustrated method 700 also includes, generating, using the electronic processor, on the user interface 315, a second graphical representation based on the metadata (at block 750). The metadata includes sufficient information to regenerate the graphical representation without the need for analyzing the underlying data needed to generate the first graphical representation. The electronic processor uses the metadata to quickly generate the second graphical representation in the second active session. As noted above, the time needed to generate the graphical representation depends on the amount of data that is to be processed. For example, the time may depend on the size of the patent profile of the entity. By storing the metadata, the graphical representation may be quickly generated. The electronic processor may continue to process the patent profile of the entity in the background while the graphical representation based on the metadata is displayed. The graphical representation may then be updated as needed after the data is processed by the electronic processor. For example, in some instances the databases 150 and 155 may be updated with new patent documents between the first active session and theAttorney Docket No.013350-0004-WO01 second active session. In these instances, the electronic processor provides the graphical representation based on the data from the first active session using the metadata stored during the first active session. In the background, the electronic processor analyzes data regarding the new patent documents and can revise the graphical representation based on the updated data on the databases 150, 155. The method 700 may be used to save and retrieve a plurality of representations. The plurality of representations may be stored in relation to the user account and a selection of an entity.

[0088] Visualizations may be generated based on a set of initial parameters. These visualizations may remain static when presented to the user. However, in some situations, it may be desirable to interact with the visualizations to refine the visualization and the data upon which the visualization is generated. One solution to allow for refining of the visualization is to have user-configurable parameters selectable using dropdown boxes, textboxes, and other user- interface tools. However, an increase in the steps needed to provide a user-defined visualization can decrease user experience.

[0089] FIG.15 is a flowchart of a method 800 for data set genome mapping and comparison that improves the user experience. The method 800 is described as being performed by the server 145, and in particular the controller 200 included in the server, for example, when executing the data server genome server application. Accordingly, the electronic processor as referred to with respect to method 800 may refer to the controller 200 and corresponding processing unit 215. In some examples, the method 800 may be distributed between the controller 200 of the server 145 and the controller(s) 300 of one or more of the client devices 105-125, and between the data server genome server application and data server genome client application. The illustrated method 800 includes generating, using the electronic processor, a graphical representation of the number of patent documents assigned to a first entity (at block 810). As discussed above with respect to method 600, a graphical representation may be generated for the first entity corresponding to the patent genome or patents by CPC codes for the first entity during an active session.

[0090] The method 800 includes receiving, using the electronic processor via the user interface 315, a selection of a section of the graphical representation (at block 820). The selected section covers a subset of patent classifications of the set of patent classifications. The selectionAttorney Docket No.013350-0004-WO01 may be received by a clicking action, a drawing action, or the like. For example, the user may individually click the classifications, draw a shape (box, oval, freeform shape, etc.) around the desired classifications within the graphical representation, and the like. The selection action may be received within the user interface on the graphical representation, or in any other manner via the user interface, such as in one or more locations separate from the graphical representation (e.g., by selecting one or more CPC Sections, Sub-Sections, Groups, Sub-Groups, or other classifications using the neighbor settings section 455 of the global settings section 460 illustrated in FIG.6), or a combination thereof.

[0091] The illustrated method 800 also includes determining, using the electronic processor, a patent profile of the entity in response to receiving the selection of the section of the graphical representation (at block 830). The patent profile is generated based on a number of patent documents assigned to the entity for the subset of patent classifications represented by the selected section. The electronic processor determines the patent profile for the entity by using only the subset of patent classifications. The electronic processor determines a patent profile by determining the number of patent documents assigned to each classification system code in the selected subset of classifications. This patent profile (e.g., second patent profile of the first entity) may be different from the patent profile (e.g., first patent profile of the first entity) determined at block 620 of method 600.

[0092] The illustrated method 800 also includes generating, using the electronic processor, on the user interface, a list of entities ranked based on similarity of patent profile with respect to the selected section of the patent profile of the entity (at block 840). The electronic processor compares the selected section of the patent profile of the entity with the patent profile of other entities, and determines the similarity of each entity with the selected entity in terms of these patent profiles. The selected section of the patent profile of the selected entity may be compared to the patent profiles of other entities that includes only the selected classifications. Specifically, the electronic processor can compare the patent profiles and generate a similarity score. The entity with the higher similarity score is placed higher in the ranked list of entities. In some examples, the comparison may be performed with respect to only the classification codes selected at block 820.Attorney Docket No.013350-0004-WO01

[0093] FIG.16 is a flowchart of a method 900 for data set genome mapping and comparison that also improves the user experience. The method 900 is described as being performed by the server 145, and in particular the controller 200 included in the server, for example, when executing the data server genome server application. Accordingly, the electronic processor as referred to with respect to method 900 may refer to the controller 200 and corresponding processing unit 215. In some examples, the method 900 may be distributed between the controller 200 of the server 145 and the controller(s) 300 of one or more of the client devices 105-125, and between the data server genome server application and data server genome client application. The illustrated method 900 includes receiving, using the electronic processor, via the user interface 315, a selection of a first entity (at block 910). The selection of the entity may be received via the search section 425. A user may search for the first entity using the search textbox 445, and select the desired entity from the search list 450 as the first entity. The method 900 includes determining, using the electronic processor, a patent profile of the first entity (at block 920). The patent profile is determined based on a number of patent documents assigned to the first entity for a plurality of patent classifications. The patent profile may be based on the patent genome of the first entity. The electronic processor determines a patent profile by determining the number of patent documents assigned to each classification system code in the patent genome.

[0094] The illustrated method 900 also includes generating, using the electronic processor, on the user interface 315, a list of entities ranked based on a similarity of patent profile with respect to the first entity (at block 930). The electronic processor compares the patent profile of the first entity with the patent profile of other entities, and determines the similarity of each entity with the first entity in terms of their patent profiles. Specifically, the electronic processor compares the patent profiles and generates a similarity score for each comparison. The entity with the higher similarity score is placed higher in the ranked list of entities.

[0095] The illustrated method 900 also includes receiving, using the electronic processor via the user interface 315, a selection action of a second entity from the list of entities (at block 940). The electronic processor may receive a user selection action (e.g., a user click) on the second entity in the Neighbors subsection 510 (see FIG.10). In one example, the electronic processor may detect a drag and drop action with respect to the first entity and the second entity. For example, the electronic processor may detect that a third graphical representation of the numberAttorney Docket No.013350-0004-WO01 of patent documents assigned to the second entity is dragged and dropped on the first graphical representation. In another example, the electronic processor may detect an entity in the list of entities being dragged onto the first graphical representation.

[0096] The illustrated method 900 also includes generating, using the electronic processor, a second graphical representation of all patent documents assigned to the first entity and the second entity in response to receiving the selection action (at block 950). The second graphical representation illustrates a number of patent documents for each patent classification for the first entity and the second entity, for example as shown in FIGs. 10 and 11. In one example, the second graphical illustration is a histogram, and includes classes of the classification system code on the X-axis and a number of patent documents on the Y-axis. The second graphical illustration may include different types of visualizations as discussed above. In some examples, as illustrated in FIG. 11, the first and second graphical illustrations are presented (i.e., simultaneously) within a user interface adjacent each other (i.e., as separate illustrations). However, in other embodiments, the first and second graphical illustrations may be provided in an (at least partially) overlapping format simultaneously (i.e., displayed at the same time as part of the user interface).

[0097] Patent databases, such as the databases 150, 155, include a large amount of data. Before visualizations can be generated, relevant data from the databases 150, 155 is loaded into local memory (e.g., memory 220). Relevant data is extracted from the databases 150, 155 using queries that provide the relevant parameters. Data satisfying the relevant parameters is then output from the database as a data frame. Based on the size of the data in the database, outputting the data frame based on the query may take a large amount of time. When a request for a new visualization or a change in current visualization is received, rerunning a query may introduce a large delay between receiving the request and updating the visualization, resulting in a poor user experience. Additionally, the data frame received from the database may not be in the right format to be used to generate the updated visualizations.

[0098] FIG.17 is a flowchart of a method 1000 for data set genome mapping and comparison that improves the user experience. FIG. 18 illustrates a process flow related to the method 1000 for data set genome mapping and comparison. The method 1000 is described as being performed by the server 145, and in particular the controller 200 included in the serverAttorney Docket No.013350-0004-WO01 145, for example, when executing the data server genome server application. Accordingly, the electronic processor as referred to with respect to method 1000 may refer to the controller 200 and corresponding processing unit 215. In some examples, the method 1000 may be distributed between the controller 200 of the server 145 and the controller(s) 300 of one or more of the client devices 105-125, and between the data server genome server application and data server genome client application.

[0099] Referring to FIGS.17 and 18, the method 1000 includes receiving, using the electronic processor, via the user interface 315, a selection of one or more of a time domain 1110, one or more patent classifications 1120, and a criteria for eligible patent parties 1130 (at block 1010). The selection of the time domain 1110, the one or more patent classifications 1120, and the criteria for eligible patent parties 1130 may be received via the settings section 430. In some examples, the time domain 1110, the one or more patent classifications 1120, and the criteria for eligible patent parties 1130 may have default values that may be changed using the settings section 430. The time domain 1110 refers to a time period having a start date and an end date, or multiple time periods having their own inclusive or exclusive start dates and end dates. The time domain 1110 relates to one or more of a priority date, a filing date, a publication date, a grant date, or the like of the patent documents. The one or more patent classifications 1120 relate to a selection of CPC codes, IPC codes, USPC codes, or the like. The one or more patent classifications 1120 may pertain to any of the hierarchical levels of the patent classification codes. For example, the one or more patent classifications 1120 may be received as a plurality of subgroups of the CPC codes. The criteria for eligible patent parties 1130 may specify the entities that can be a patent party. For example, the criteria for eligible patent parties 1130 may include one or more assignees, inventive entities, applicants, law firms, or the like, and any combination of such parties. In some examples, a default value may be selected for each of the time domain 1110, the one or more patent classifications 1120, and the criteria for eligible patent parties 1130. For example, the default value for the time domain 1110 may include the last 10 years, the last 20 years, the last 30 years, all of the time period for which data is available in the patent database, or the like. The default value for the one or more patent classifications 1120 may include all classifications of a particular classification system (e.g., all classifications of the CPC system). The default value for the criteria for eligible patent parties 1130 may include all patent parties. Thus, the electronic processor may be configured to receive user-based “selections,”Attorney Docket No.013350-0004-WO01 default “selections,” or a combination thereof (e.g., depending on whether a user has adjusted one or more of the default values).

[0100] The illustrated method 1000 also includes querying, using the electronic processor, a patent database 1140 (at block 1020). The query may be generated based on selected values of one or more of the time domain 1110, the one or more patent classifications 1120, and the criteria for eligible patent parties 1130. The patent database 1140 is, for example, any one of the databases 150, 155. In some implementations, the electronic processor automatically generates a query based on the selections received for the time domain 1110, the one or more patent classifications 1120, and the criteria for eligible patent parties 1130. The query may be generated in any appropriate querying language, for example, SQL, LINQ, or the like. The electronic processor then provides the query to the patent database 1140 (e.g., directly or via one or more intermediary devices) to retrieve the relevant data from the patent database 1140.

[0101] The illustrated method 1000 also includes receiving, using the electronic processor, a long-form data frame 1150 in response to querying the patent database 1140 (at block 1030). The long-form data frame 1150 is received from the patent database 1140 in response to the query generated by the electronic processor based on the selection of the time domain 1110, the one or more patent classifications 1120, and the criteria for eligible patent parties 1130. The long-form data frame 1150 may be received in extensible markup language (XML), comma separated values (CSV), or other similar data formats. FIG.19 illustrates an example of the long-form data frame 1150. The long-form data frame 1150 includes data based on, for example, a party identifier 1210, a patent classification 1220, and a number of patents 1230. A separate entry may be included for each combination of the party identifier 1210 and the patent classification 1220. As described above, the patent classification 1220 may refer to any hierarchical level of the patent classification codes. The number of patents 1230 refers to the number of patents for the particular patent party in the particular patent classification 1220. The long-form data frame 1150 may ignore zero value data such that when a patent party does not have any patents in a particular patent classification 1220, a separate row of the long-form data frame 1150 is not generated with such a zero value. That is, only patent parties and patent classifications 1120 that include non-zero number of patents 1230 may be represented in the long-form data frame 1150, which may reduce a size of the long-form data frame 1150.Attorney Docket No.013350-0004-WO01

[0102] Referring back to FIGS.17 and 18, the method 1000 also includes converting, using the electronic processor, the long-form data frame 1150 into a wide-form data frame 1160 by pivoting the long-form data frame 1150 (at block 1040). FIG. 20 illustrates an example of the wide-form data frame 1160. The electronic processor pivots, for example, a table of the long- form data frame 1150 such that each party identifier 1210 is provided on a separate single row and each patent classification 1220 is provided on a separate single column (or vice versa). The number of patents 1230 can then be provided in a cell corresponding to the party identifier 1210 and the patent classification 1220. The electronic processor may fill in zeroes for cells that do not have any data (i.e., number of patents 1230). The wide-form data frame 1160 may therefore take the form of an n row and c column data table, where n is the number of patent parties and c is the number of patent classifications (or vice versa).

[0103] The illustrated method 1000 also includes storing, using the electronic processor, the wide-form data frame 1160 in a local memory (at block 1050). The local memory is, for example, any of the memories 220, 330. The wide-form data frame 1160 allows for the additional operations as described herein to be performed on the data to generate the desired visualizations. As described above, interacting directly with the database to generate the visualizations may be require a large amount of processing time that degrades the user experience. Generating the wide-form data frame 1160 allows for faster processing such that the desired visualizations may be generated more quickly – including updates to visualizations. Storing the wide-form data frame 1160 in local memory also allows for quick access and processing, which further increases the speed of generating visualizations.

[0104] Visualizations can be easily generated from the wide-form data frame 1160. An example of generating a visualization of a Patent Genome as displayed in the Patent Genome subsection 500 is explained below with respect to the wide-form data frame 1160. The visualization of the Patent Genome may be generated by representing the patent classifications on a first axis (e.g., one of the X-axis and the Y-axis) and representing the number of patent documents on a second axis (e.g., the other of the X-axis and the Y-axis). The chart element (e.g., histogram, bar, point, bubble, etc.) is provided for each patent classification representing the number of patent documents for that patent classification belonging to the selected patent party. Other visualizations as described above may be similarly generated.Attorney Docket No.013350-0004-WO01

[0105] In some examples, the method 1000 may also optionally include generating, using the electronic processor, a data matrix 1170 by performing a data transformation on the wide-form data frame 1160. FIG.21 illustrates an example of the data matrix 1170. The electronic processor may perform any one or more of the following: filtering for near zero variance, centering / scaling, logarithmic / power transformations, dimension reduction, or the like on the wide-form data frame 1160 to generate the data matrix 1170. The resulting data matrix 1170 includes a similar tabular structure as the wide-form data frame 1160, but with the number of patents changed to numeric features 1240 generated based on the data transformations performed on the wide-form data frame 1160. In the resulting data matrix 1170, each patent classification may be considered as a dimension (or direction). Each party includes a numeric feature for each of the patent classifications, which can be considered as the magnitude for that dimension or direction. Each party therefore has a vector defined by the magnitudes and dimensions (i.e., numeric features and patent classifications). That is, the data matrix 1170 defines coordinates in a multi-dimensional coordinate system for each of the plurality of patent parties. This allows the electronic processor to determine the distance between each party. Similar to the wide-form data frame 1160, the electronic processor may also store the data matrix 1170 in a local memory, albeit, in a different location than the wide-form data frame 1160.

[0106] Referring to FIG.22, the data matrix 1170 can be used to determine the nearest neighbors of a patent party. The electronic processor may receive a selection of or use default values of a patent party 1180 and a distance metric 1190 (e.g., Euclidean, Manhattan, Chebyshev, and Angular, such as those shown in FIG.10). The selections of the patent party 1180, the distance metric 1190, the time domain 1110, the one or more patent classifications 1120, and the criteria for eligible patent parties 1130 may be received as part of a single step (e.g., a single input) or may be received as part of multiple steps. In one example, a selection of the patent party 1180 may be received via the search sub-pane 440, and the remaining settings may be received via the settings section 430. The remaining settings may be default setting that are pre-loaded in the setting section 430. The patent party 1180 is one from a plurality of patent parties represented in the data matrix 1170. As indicated above, the distance metric 1190 is, for example, one of the several methodologies for calculating distance (e.g., Euclidean, Manhattan, Angular, Chebyshev, etc.). The electronic processor determines the distance between the patent party 1180 and some or all of the plurality of patent parties represented in the data matrix 1170Attorney Docket No.013350-0004-WO01 based on the distance metric 1190. Distance metrics 1190 express the similarity / dissimilarity between samples of a data set that are each represented as a point in an n-dimensional space. As explained above, the patent classifications may define the different dimensions of the n- dimensional space. Similarity measures are typically numerical representations of how close two samples appear in the n-dimensional space, with a higher score indicating a closer relationship. Some well-known distance metrics commonly used to find the distance between elements of a metric set are Euclidean, Manhattan, Hamming, Canberra, Angular, and Chebyshev. These metrics differ in how they measure and represent distances, with certain metrics being more applicable in different applications and settings.

[0107] Euclidean distance refers to the direct distance between two points, or the shortestdistance between two vectors. Given two vectors v and w, Euclidean distance is:

[0108] Manhattan distance is the sum of the absolute differences between two vectors:

[0109] Angular distance does not directly evaluate the length of the vectors, but instead considers the cosine of the angle formed between them as projected from the same origin. The smaller the angle between them, the more similar the two vectors are in direction:

[0110] The distance metrics described herein are for exemplary purposes only. Any distance metric, including custom metrics having weighted dimensions, may be used to determine the nearest neighbors. After the distances are determined, the electronic processor may generate a list of neighbors 1200 by sorting the distances in ascending order such that the nearest neighbors are at the top of the list. This list of nearest neighbors 1200, or a portion thereof, is displayed in the nearest neighbor section 510 (as shown in FIGS.9 and 10).Attorney Docket No.013350-0004-WO01

[0111] The storing of the wide-form data frame 1160 and / or the data matrix 1170 allows for faster processing of the data. For example, when a second patent party and / or a second distance matrix is selected, the electronic processor may quickly regenerate the ranked list of nearest neighbors 1200 from the wide-form data frame 1160 and / or the data matrix 1170. That is, a new query need not be submitted to the patent database, which would result in a delay in generating the new ranked list or visualizations.

[0112] Referring to FIG.23, the data matrix 1170 can also or instead be used to generate a neighbor cluster. With reference also to FIG. 12, the electronic processor may receive a selection of a number of neighbors in the cluster 1250, a clustering method 1260, and a number of clusters 1270. The selection of these three parameters may be received, for example, via the neighbors slider 550, the clusters slider 555, and the clustering method selector 560, respectively. The number of neighbors in the cluster 1250 selection is used to limit the number of neighbors represented in the resulting cluster visualization. The clustering method 1260 is, for example, one of the several linkage methodologies for clustering (e.g., Ward’s method, complete, average, etc.). The electronic processor may determine the nearest neighbors using the data matrix 1170 as described above, and filter the nearest neighbors to limit the list of nearest neighbors to the selected number of neighbors in the cluster 1270. For example, when the selected number of neighbors in the cluster 1270 is 40, the electronic processor keeps only the 40 nearest neighbors (i.e., top 40 closest to the selected patent party) from the list of neighbors.

[0113] The electronic processor may determine the angular distance between all parties in the filtered data matrix 1280 to generate a distance matrix 1290. The distance matrix 1290 includes an N x N matrix, where N is the selected number of neighbors in the cluster. The distance matrix 1290 therefore includes the calculated pairwise distance between neighbors. The electronic processor performs a hierarchical clustering methodology on the distance matrix 1290 to assign each party to a particular cluster label, and generate a neighbor cluster 1295. The number of cluster labels is the same as the selected number of clusters. Hierarchical clustering includes assigning labels iteratively such that each patent party is assigned to its own label. After each iteration, clusters that are closest to each other in distance are then joined together to form a new cluster, proceeding until all the parties have been assigned to the selected number of clusters. The Dendrogram plot, as provided in the Dendrogram subsection 565 of FIG.12,Attorney Docket No.013350-0004-WO01 illustrates the distances between neighbors and the number of clusters that were selected, highlighting the level at which clusters (i.e., cluster labels) are assigned to patent parties.

[0114] The clustering method 1260 may define how the patent parties are linked together. For example, different linkage methods may be used to form the clusters in the hierarchical clustering methodology described above. A ‘Single’ linkage method merges clusters based on the shortest distance between any two clusters. A ‘Complete’ linkage method merges clusters based on the maximum distance between any two clusters. An ‘Average’ linkage method merges two clusters that have the smallest average distance between them. A ‘Centroid’ linkage method merges the closest pair of clusters based on the distance between their centroids (i.e., the average of all data points in each cluster). A ‘Ward’s’ method minimizes the within-cluster variance, and merges the pair of clusters that results in the lowest increase in variance. As noted above, the clusters are merged together until the selected number of clusters are formed.

[0115] FIG.24 is a flowchart of a method 1300 for data set genome mapping and comparison. The method 1300 is described as being performed by the server 145, and in particular the controller 200 included in the server 145, for example, when executing the data server genome server application. Accordingly, the electronic processor as referred to with respect to method 1000 may refer to the controller 200 and corresponding processing unit 215. In some examples, the method 1300 may be distributed between the controller 200 of the server 145 and the controller(s) 300 of one or more of the client devices 105-125, and between the data server genome server application and data server genome client application.

[0116] The method 1300 includes querying, using the electronic processor, a patent database 1140 (at block 1310). As discussed above, the query may be generated based on selected values or default values of one or more of the time domain 1110, the one or more patent classifications 1120, and the criteria for eligible patent parties 1130. The electronic processor provides the query to the patent database 1140 (e.g., directly or via one or more intermediary devices).

[0117] The illustrated method 1300 also includes receiving, using the electronic processor, a data frame in response to querying the patent database 1140 (at block 1320). The data frame includes a patent profile for each of a plurality of patent parties, wherein the patent profile for each of the plurality of patent parties includes a number of patent documents associated with theAttorney Docket No.013350-0004-WO01 each of the plurality of patent parties in each of a plurality of patent classifications. The data frame is, for example, the long-form data frame 1150 as shown in FIG.19.

[0118] The illustrated method 1300 also includes receiving, using the electronic processor, on a user interface 315, a selection of a patent party from the plurality of patent parties (at block 1330). The selection of the patent party may be received via the search section 425. In response to receiving the selection of the patent party, the method also includes generating, using the electronic processor, on the user interface 315, a graphical representation illustrating the patent profile of the patent party from the data frame (at block 1340). As discussed above, the electronic processor may convert the long-form data frame 1150 into the wide-form data frame 1160 and / or the data matrix 1170. The graphical representation is, for example, a visualization of the patent profile of the patent party generated by representing the patent classifications on a first axis (e.g., one of the X-axis or the Y-axis) and representing the number of patent documents on a second axis (e.g., the other of the X-axis or the Y-axis). The chart element (e.g., histogram, bar, point, bubble, etc.) is provided for each patent classification, and represents the number of patent documents included in respective patent classifications associated with the selected patent party.

[0119] The illustrated method 1300 also includes determining, using the electronic processor, a plurality of distances based on a distance metric, each of the plurality of distances being a distance between the patent party and at least a subset of the other patent parties included in the plurality of patent parties represented in the data frame 1160 (at block 1350). The method 1300 also includes generating, using the electronic processor, on the user interface, a ranked list of the plurality of patent parties based on the plurality of distances (at block 1360). The electronic processor may determine the plurality of distances and display the ranked list similarly as described above with respect to FIG.22.

[0120] By creating the wide-form data frame 1160, a user can change the settings associated with a visualization without re-running the query sent to the patents database or the data frame that is stored locally with summarized class information for parties (e.g., summarized CPC codes). Similarly, a change to the methodology for neighbors only affects the portions of the method 1000 where neighbors are calculated, so only these parts of the method 1000 need to be re-run and presented within the user interface. Implementation of reactive expressions in this way allows the methodology to be run in a manner that is computationally feasible. In otherAttorney Docket No.013350-0004-WO01 words, the system and applications described herein can implement some or all aspects of the methodology described above to allow a user to find patent neighbors on command. That is, the application allows the user to update / change settings and find patent neighbors given their selection. For example, rather than running the analysis for neighbors in advance and defining a set of tables from which a user queries, the application allows the user to run the methodology at their own direction in real-time. To accomplish this, the application makes use of reactive programming to allow an end user to define the inputs and update and change the resulting outputs. In particular, steps in the middle of the method 1000 (between the selections or inputs and the output) are intermediate pieces used in delivering outputs and, in particular, these intermediate pieces are reactive expressions in that they are defined by the configurable settings (e.g., cached locally in memory), which are only updated if a change is made.

[0121] As compared to traditional programming based on a sequential model (code executed in a pre-determined, step-by-step order), this reactive programming is based on a model of data streams and asynchronous processing. In such reactive programming, the program processes a continuous stream of events or data changes, and the processing is done in a reactive, as-needed manner. Thus, as compared to traditional programming, this reactive programming is more efficient when processing large amounts of data, as it avoids the overhead of having to loop through data structures. Additionally, this reactive programming can be used to build more responsive user interfaces, as it can update the user interface instantly when changes occur.

[0122] Thus, the disclosure provides, among other things, data set genome mapping and comparison systems and methods for faster processing of visualizations and improved user experience. The embodiments described above and illustrated in the figures are presented by way of example only and are not intended as a limitation upon the concepts and principles of the present invention. As such, it will be appreciated by one having ordinary skill in the art that various changes in the elements and their configuration and arrangement are possible without departing from the spirit and scope of the present invention.

[0123] Implementations of the present disclosure are disclosed in the following example clauses:

[0124] Clause 1. A method comprising: querying, using an electronic processor, a patent database; receiving, using the electronic processor, a data frame in response to the query, the dataAttorney Docket No.013350-0004-WO01 frame including a patent profile for each of a plurality of patent parties, wherein the patent profile for each of the plurality of patent parties includes a number of patent documents associated with the each of the plurality of patent parties in each of a plurality of patent classifications; receiving, using the electronic processor, via a user interface, a selection of a patent party from the plurality of patent parties; and in response to receiving the selection generating, using the electronic processor, on the user interface, a graphical representation illustrating the patent profile of the patent party from the data frame, determining, using the electronic processor, a plurality of distances based on a distance metric, each of the plurality of distances being a distance between the patent party and each other one of the plurality of patent parties represented in the data frame, and generating, using the electronic processor, on the user interface, a ranked list of the plurality of patent parties based on the plurality of distances.

[0125] Clause 2. The method of clause 1, wherein the graphical representation includes a histogram representing the number of patent documents associated with the patent party for each of the plurality of patent classifications.

[0126] Clause 3. The method of any one of clauses 1-2, wherein the patent party is a first patent party and the graphical representation is a first graphical representation, the method further comprising: receiving, using the electronic processor, via the user interface, a selection of a second patent party from the ranked list of the plurality of patent parties; and generating, using the electronic processor, on the user interface, a second graphical representation illustrating the patent profile of the patent party simultaneously with the first graphical representation in response to receiving the selection of the second patent party.

[0127] Clause 4. The method of any one of clauses 1-3, further comprising: generating, using the electronic processor, a data matrix by performing a data transformation on the data frame, wherein the plurality of distances is determined from the data matrix.

[0128] Clause 5. The method of any one of clauses 1-4, wherein the distance metric is a first distance metric, the plurality of distances is a first plurality of distances, and the ranked list is a first ranked list, further comprising: receiving, using the electronic processor, via the user interface, a selection of a second distance metric; determining, using the electronic processor, a second plurality of distances based on a second distance metric, each of the second plurality of distances being a distance between the patent party and each other one of the plurality of patentAttorney Docket No.013350-0004-WO01 parties represented in the data frame; and generating, using the electronic processor, on the user interface, a second ranked list of the plurality of patent parties based on the second plurality of distances.

[0129] Clause 6. A method comprising: querying, using an electronic processor, a patent database based on a time domain and one or more patent classifications; receiving, using the electronic processor, a long-form data frame in response to querying the patent database; converting, using the electronic processor, the long-form data frame into a wide-form data frame by pivoting the long-form data frame; and storing, using the electronic processor, the wide-form data frame in a local memory.

[0130] Clause 7. The method of clause 6, further comprising: generating, using the electronic processor, a data matrix by performing a data transformation on the wide-form data frame; and storing, using the electronic processor, the data matrix in the local memory.

[0131] Clause 8. The method of clause 7, wherein the data matrix defines coordinates for each of a plurality of patent parties represented in the data matrix, the method further comprising: determining, using the electronic processor, a plurality of distances, each of the plurality of distances being a distance between a patent party and one of a plurality of patent parties represented in the data matrix determined based on a distance metric; and generating, using the electronic processor, on a user interface, a ranked list of the plurality of patent parties based on the plurality of distances.

[0132] Clause 9. The method of any one of clauses 6-8, wherein the patent party is a first patent party, the plurality of distances is a first plurality of distances, and the ranked list is a first ranked list, the method further comprising: receiving, using the electronic processor, a selection of a second patent party; determining, using the electronic processor, a second plurality of distances, each of the second plurality of distances being a distance between the second patent party and one of the plurality of patent parties in the data matrix determined based on the distance metric; and generating, using the electronic processor, on the user interface, a second ranked list of the plurality of patent parties based on the second plurality of distances.

[0133] Clause 10. The method of any one of clauses 6-9, wherein the distance metric is a first distance metric, the plurality of distances is a first plurality of distances, and the ranked list is a first ranked list, the method further comprising: receiving, using the electronic processor, aAttorney Docket No.013350-0004-WO01 selection of a second distance metric; determining, using the electronic processor, a second plurality of distances, each of the second plurality of distances being a distance between the patent party and one of the plurality of patent parties in the data matrix determined based on the second distance metric; and generating, using the electronic processor, on the user interface, a second ranked list of the plurality of patent parties based on the second plurality of distances.

[0134] Clause 11. A method comprising: generating, using an electronic processor, on a user interface, a first graphical representation of a number of patent documents assigned to an entity during a first active session, the first graphical representation illustrating the number of patent documents assigned to the entity in a plurality of patent classifications; storing, using the electronic processor, metadata corresponding to the first graphical representation in relation with the entity; receiving, using the electronic processor, a selection of the entity during a second active session; retrieving, using the electronic processor, the metadata stored during the first active session in response to receiving the selection of the entity during the second active session; and generating, using the electronic processor, on the user interface, a second graphical representation based on the metadata.

[0135] Clause 12. The method of clause 11, further comprising: determining, using the electronic processor, a patent profile of the entity based on a number of patent documents assigned to the entity for the plurality of patent classifications.

[0136] Clause 13. The method of any one of clauses 11-12, wherein the entity is a first entity, further comprising: generating, using the electronic processor, on the user interface, a list of entities ranked based on a similarity of a patent profile of entities with respect to the first entity.

[0137] Clause 14. The method of any one of clauses 11-13, further comprising: generating, using the electronic processor, on the user interface, a third graphical representation of a number of patent documents assigned to a second entity from the list of entities, the third graphical representation illustrating a number of patent documents for the plurality of patent classifications for the second entity.

[0138] Clause 15. The method of clause 14, wherein the first graphical representation and the third graphical representation illustrate a patent genome of the first entity and the second entity.Attorney Docket No.013350-0004-WO01

[0139] Clause 16. The method of clause 14, wherein the metadata is a first metadata, further comprising: storing, using the electronic processor, second metadata corresponding to the third graphical representation in relation with the entity during the first active session; and retrieving, using the electronic processor, the second metadata stored during the first active session in response to receiving the selection of the first entity during the second active session.

[0140] Clause 17. The method of any one of clauses 11-16, wherein the patent profile is a first patent profile, further comprising: receiving, using the electronic processor via the user interface, a selection of a section of the first graphical representation, the section covering a subset of patent classifications of the plurality of patent classifications; generating, using the electronic processor, a second patent profile of the first entity in response to receiving the selection of the section of the first graphical representation, the second patent profile generated based on a number of patent documents assigned to the first entity for the subset of patent classifications; and generating, using the electronic processor, on the user interface, a second list of entities ranked based on similarity of patent profile with respect to the second patent profile.

[0141] Clause 18. A method comprising: generating, using an electronic processor, on a user interface, a graphical representation of a number of patent documents assigned to an entity, the graphical representation illustrating a number of patent documents for a plurality of patent classifications for the entity; receiving, using the electronic processor via the user interface, a selection of a section of the graphical representation, the section covering a subset of patent classifications of the plurality of patent classifications; determining, using the electronic processor, a patent profile of the entity in response to receiving the selection of the section of the graphical representation, the patent profile generated based on a number of patent documents assigned to the entity for the subset of patent classifications; and generating, using the electronic processor, on the user interface, a list of entities ranked based on similarity of patent profile with respect to the patent profile of the entity.

[0142] Clause 19. The method of clause 18, wherein the patent profile is a first patent profile, further comprising: determining, using the electronic processor, a second patent profile of the entity based on the plurality of patent classifications.

[0143] Clause 20. The method of clause 19, wherein the entity is a first entity and the list of entities is a first list of entities, further comprising: generating, using the electronic processor, onAttorney Docket No.013350-0004-WO01 the user interface, a second list of entities ranked based on a similarity of patent profile of entities with respect to the second patent profile.

[0144] Clause 21. The method of any one of clauses 18-20, wherein the graphical representation is a first graphical representation, further comprising: generating, using the electronic processor, on the user interface, a second graphical representation of a number of patent documents assigned to the first entity for the plurality of patent classifications.

[0145] Clause 22. The method of any one of clauses 18-21, further comprising: generating, using the electronic processor, on the user interface, a third graphical representation of a number of patent documents assigned to a second entity from the list of entities, the third graphical representation illustrating a number of patent documents for the plurality of patent classifications for the second entity.

[0146] Clause 23. The method of clause 22, wherein the second graphical representation and the third graphical representation illustrate a patent genome of the first entity and the second entity.

[0147] Clause 24. The method of any one of clauses 18-23, wherein the graphical representation is a first graphical representation, further comprising: generating, using the electronic processor, on the user interface, a second graphical representation of a number of patent documents assigned to a second entity from the list of entities, the second graphical representation illustrating a number of patent documents for the subset of patent classifications for the second entity.

[0148] Clause 25. The method of any one of clauses 18-24, wherein the graphical representation is a first graphical representation and the entity is a first entity, further comprising: receiving, using the electronic processor via the user interface, a selection action of a second entity from the list of entities, the selection action received with respect to the first entity; and generating, using the electronic processor, a second graphical representation of a number patent documents assigned to the first entity and the second entity in response to receiving the selection action, the second graphical representation illustrating the number of patent documents for the subset of patent classifications for the first entity and the second entity.Attorney Docket No.013350-0004-WO01

[0149] Clause 26. A method comprising: receiving, using an electronic processor via a user interface, a selection of a first entity; determining, using the electronic processor, a patent profile of the first entity based on a number of patent documents assigned to the first entity for a plurality of patent classifications; generating, using the electronic processor, on the user interface, a list of entities ranked based on a similarity of a patent profile with respect to the patent profile of the first entity; receiving, using the electronic processor, via the user interface, a selection action of a second entity from the list of entities, the selection action received with respect to the first entity; and generating, using the electronic processor, a graphical representation of a number of patent documents assigned to the first entity and a number of patent documents assigned to the second entity in response to receiving the selection action, the graphical representation illustrating the number of patent documents assigned to the first entity and the number of patent documents assigned to the second entity for the plurality of patent classifications.

[0150] Clause 27. The method of clause 26, wherein the graphical representation is a first graphical representation, further comprising: generating, using the electronic processor, on the user interface, a second graphical representation of a number of patent documents assigned to the first entity, the graphical representation illustrating the number of patent documents assigned to the first entity in a plurality of patent classifications, wherein the selection action is received with respect to the second graphical representation.

[0151] Clause 28. The method of claim of clause 27, further comprising: storing, using the electronic processor, metadata corresponding to the first graphical representation and the second graphical representation in relation with the first entity during a first active session; receiving, using the electronic processor, the selection of the first entity during a second active session; and retrieving, using the electronic processor, the metadata stored during the first active session in response to receiving the selection of the first entity during the second active session.

[0152] Clause 29. The method of any one of clauses 26-28, further comprising: storing, using the electronic processor, metadata corresponding to the graphical representation in relation with the first entity during a first active session; receiving, using the electronic processor, the selection of the first entity during a second active session; and retrieving, using the electronic processor,Attorney Docket No.013350-0004-WO01 the metadata stored during the first active session in response to receiving the selection of the first entity during the second active session.

[0153] Clause 30. The method of any one of clauses 26-29, wherein the graphical representation illustrates a patent genome of the first entity and the second entity.

[0154] Various features and advantages are set forth in the following claims.

Claims

Attorney Docket No.013350-0004-WO01 CLAIMS What is claimed is:

1. A method comprising: querying, using an electronic processor, a patent database; receiving, using the electronic processor, a data frame in response to the query, the data frame including a patent profile for each of a plurality of patent parties, wherein the patent profile for each of the plurality of patent parties includes a number of patent documents associated with the each of the plurality of patent parties in each of a plurality of patent classifications; receiving, using the electronic processor, via a user interface, a selection of a patent party from the plurality of patent parties; and in response to receiving the selection generating, using the electronic processor, on the user interface, a graphical representation illustrating the patent profile of the patent party from the data frame, determining, using the electronic processor, a plurality of distances based on a distance metric, each of the plurality of distances being a distance between the patent party and each other one of the plurality of patent parties represented in the data frame, and generating, using the electronic processor, on the user interface, a ranked list of the plurality of patent parties based on the plurality of distances.

2. The method of claim 1, wherein the graphical representation includes a histogram representing the number of patent documents associated with the patent party for each of the plurality of patent classifications.

3. The method of claim 2, wherein the patent party is a first patent party and the graphical representation is a first graphical representation, the method further comprising: receiving, using the electronic processor, via the user interface, a selection of a second patent party from the ranked list of the plurality of patent parties; andAttorney Docket No.013350-0004-WO01 generating, using the electronic processor, on the user interface, a second graphical representation illustrating the patent profile of the patent party simultaneously with the first graphical representation in response to receiving the selection of the second patent party.

4. The method of claim 2, further comprising: generating, using the electronic processor, a data matrix by performing a data transformation on the data frame, wherein the plurality of distances is determined from the data matrix.

5. The method of claim 3, wherein the distance metric is a first distance metric, the plurality of distances is a first plurality of distances, and the ranked list is a first ranked list, further comprising: receiving, using the electronic processor, via the user interface, a selection of a second distance metric; determining, using the electronic processor, a second plurality of distances based on a second distance metric, each of the second plurality of distances being a distance between the patent party and each other one of the plurality of patent parties represented in the data frame; and generating, using the electronic processor, on the user interface, a second ranked list of the plurality of patent parties based on the second plurality of distances.

6. A method comprising: querying, using an electronic processor, a patent database based on a time domain and one or more patent classifications; receiving, using the electronic processor, a long-form data frame in response to querying the patent database; converting, using the electronic processor, the long-form data frame into a wide-form data frame by pivoting the long-form data frame; and storing, using the electronic processor, the wide-form data frame in a local memory.

7. The method of claim 6, further comprising:Attorney Docket No.013350-0004-WO01 generating, using the electronic processor, a data matrix by performing a data transformation on the wide-form data frame; and storing, using the electronic processor, the data matrix in the local memory.

8. The method of claim 7, wherein the data matrix defines coordinates for each of a plurality of patent parties represented in the data matrix, the method further comprising: determining, using the electronic processor, a plurality of distances, each of the plurality of distances being a distance between a patent party and one of a plurality of patent parties represented in the data matrix determined based on a distance metric; and generating, using the electronic processor, on a user interface, a ranked list of the plurality of patent parties based on the plurality of distances.

9. The method of claim 8, wherein the patent party is a first patent party, the plurality of distances is a first plurality of distances, and the ranked list is a first ranked list, the method further comprising: receiving, using the electronic processor, a selection of a second patent party; determining, using the electronic processor, a second plurality of distances, each of the second plurality of distances being a distance between the second patent party and one of the plurality of patent parties in the data matrix determined based on the distance metric; and generating, using the electronic processor, on the user interface, a second ranked list of the plurality of patent parties based on the second plurality of distances.

10. The method of claim 8, wherein the distance metric is a first distance metric, the plurality of distances is a first plurality of distances, and the ranked list is a first ranked list, the method further comprising: receiving, using the electronic processor, a selection of a second distance metric; determining, using the electronic processor, a second plurality of distances, each of the second plurality of distances being a distance between the patent party and one of the plurality of patent parties in the data matrix determined based on the second distance metric; and generating, using the electronic processor, on the user interface, a second ranked list of the plurality of patent parties based on the second plurality of distances.Attorney Docket No.013350-0004-WO01 11. A method comprising: generating, using an electronic processor, on a user interface, a first graphical representation of a number of patent documents assigned to an entity during a first active session, the first graphical representation illustrating the number of patent documents assigned to the entity in a plurality of patent classifications; storing, using the electronic processor, metadata corresponding to the first graphical representation in relation with the entity; receiving, using the electronic processor, a selection of the entity during a second active session; retrieving, using the electronic processor, the metadata stored during the first active session in response to receiving the selection of the entity during the second active session; and generating, using the electronic processor, on the user interface, a second graphical representation based on the metadata.

12. The method of claim 11, further comprising: determining, using the electronic processor, a patent profile of the entity based on a number of patent documents assigned to the entity for the plurality of patent classifications.

13. The method of claim 12, wherein the entity is a first entity, further comprising: generating, using the electronic processor, on the user interface, a list of entities ranked based on a similarity of a patent profile of entities with respect to the first entity.

14. The method of claim 13, further comprising: generating, using the electronic processor, on the user interface, a third graphical representation of a number of patent documents assigned to a second entity from the list of entities, the third graphical representation illustrating a number of patent documents for the plurality of patent classifications for the second entity.

15. The method of claim 14, wherein the first graphical representation and the third graphical representation illustrate a patent genome of the first entity and the second entity.Attorney Docket No.013350-0004-WO01 16. The method of claim 14, wherein the metadata is a first metadata, further comprising: storing, using the electronic processor, second metadata corresponding to the third graphical representation in relation with the entity during the first active session; and retrieving, using the electronic processor, the second metadata stored during the first active session in response to receiving the selection of the first entity during the second active session.

17. The method of claim 14, wherein the patent profile is a first patent profile, further comprising: receiving, using the electronic processor via the user interface, a selection of a section of the first graphical representation, the section covering a subset of patent classifications of the plurality of patent classifications; generating, using the electronic processor, a second patent profile of the first entity in response to receiving the selection of the section of the first graphical representation, the second patent profile generated based on a number of patent documents assigned to the first entity for the subset of patent classifications; and generating, using the electronic processor, on the user interface, a second list of entities ranked based on similarity of patent profile with respect to the second patent profile.Attorney Docket No.013350-0004-WO01 18. A method comprising: generating, using an electronic processor, on a user interface, a graphical representation of a number of patent documents assigned to an entity, the graphical representation illustrating a number of patent documents for a plurality of patent classifications for the entity; receiving, using the electronic processor via the user interface, a selection of a section of the graphical representation, the section covering a subset of patent classifications of the plurality of patent classifications; determining, using the electronic processor, a patent profile of the entity in response to receiving the selection of the section of the graphical representation, the patent profile generated based on a number of patent documents assigned to the entity for the subset of patent classifications; and generating, using the electronic processor, on the user interface, a list of entities ranked based on similarity of patent profile with respect to the patent profile of the entity.

19. The method of claim 18, wherein the patent profile is a first patent profile, further comprising: determining, using the electronic processor, a second patent profile of the entity based on the plurality of patent classifications.

20. The method of claim 19, wherein the entity is a first entity and the list of entities is a first list of entities, further comprising: generating, using the electronic processor, on the user interface, a second list of entities ranked based on a similarity of patent profile of entities with respect to the second patent profile.

21. The method of claim 20, wherein the graphical representation is a first graphical representation, further comprising: generating, using the electronic processor, on the user interface, a second graphical representation of a number of patent documents assigned to the first entity for the plurality of patent classifications.Attorney Docket No.013350-0004-WO01 22. The method of claim 21, further comprising: generating, using the electronic processor, on the user interface, a third graphical representation of a number of patent documents assigned to a second entity from the list of entities, the third graphical representation illustrating a number of patent documents for the plurality of patent classifications for the second entity.

23. The method of claim 22, wherein the second graphical representation and the third graphical representation illustrate a patent genome of the first entity and the second entity.

24. The method of claim 18, wherein the graphical representation is a first graphical representation, further comprising: generating, using the electronic processor, on the user interface, a second graphical representation of a number of patent documents assigned to a second entity from the list of entities, the second graphical representation illustrating a number of patent documents for the subset of patent classifications for the second entity.

25. The method of claim 18, wherein the graphical representation is a first graphical representation and the entity is a first entity, further comprising: receiving, using the electronic processor via the user interface, a selection action of a second entity from the list of entities, the selection action received with respect to the first entity; and generating, using the electronic processor, a second graphical representation of a number patent documents assigned to the first entity and the second entity in response to receiving the selection action, the second graphical representation illustrating the number of patent documents for the subset of patent classifications for the first entity and the second entity.Attorney Docket No.013350-0004-WO01 26. A method comprising: receiving, using an electronic processor via a user interface, a selection of a first entity; determining, using the electronic processor, a patent profile of the first entity based on a number of patent documents assigned to the first entity for a plurality of patent classifications; generating, using the electronic processor, on the user interface, a list of entities ranked based on a similarity of a patent profile with respect to the patent profile of the first entity; receiving, using the electronic processor, via the user interface, a selection action of a second entity from the list of entities, the selection action received with respect to the first entity; and generating, using the electronic processor, a graphical representation of a number of patent documents assigned to the first entity and a number of patent documents assigned to the second entity in response to receiving the selection action, the graphical representation illustrating the number of patent documents assigned to the first entity and the number of patent documents assigned to the second entity for the plurality of patent classifications.

27. The method of claim 26, wherein the graphical representation is a first graphical representation, further comprising: generating, using the electronic processor, on the user interface, a second graphical representation of a number of patent documents assigned to the first entity, the graphical representation illustrating the number of patent documents assigned to the first entity in a plurality of patent classifications, wherein the selection action is received with respect to the second graphical representation.

28. The method of claim of claim 27, further comprising: storing, using the electronic processor, metadata corresponding to the first graphical representation and the second graphical representation in relation with the first entity during a first active session; receiving, using the electronic processor, the selection of the first entity during a second active session; and retrieving, using the electronic processor, the metadata stored during the first active session in response to receiving the selection of the first entity during the second active session.Attorney Docket No.013350-0004-WO01 29. The method of claim 26, further comprising: storing, using the electronic processor, metadata corresponding to the graphical representation in relation with the first entity during a first active session; receiving, using the electronic processor, the selection of the first entity during a second active session; and retrieving, using the electronic processor, the metadata stored during the first active session in response to receiving the selection of the first entity during the second active session.

30. The method of claim 26, wherein the graphical representation illustrates a patent genome of the first entity and the second entity.

Citation Information

Patent Citations

  • System and method delivering remotely stored applications and information

    US20180007171A1

  • Method and system for multistage candidate ranking

    US20210097472A1

  • Analytics generation for patent portfolio management

    US20230018572A1

  • Learning embedded representation of a correlation matrix to a network with machine learning

    US20240012997A1