October 30, 2024
In our pursuit of understanding the world, we continuously evaluate the credibility and relevance of various information sources and different viewpoints. When confronted with conflicting or confusing data, we draw on our experiences and beliefs to interpret and integrate the facts, forming a coherent understanding. By combining facts with new observations and applying logical reasoning, we gain a more comprehensive perspective.
What is Knowledge Synergy?
This method can be referred to as knowledge synergy, and it allows us to identify mis- and disinformation, fraud, as well as make more accurate predictions and anticipate potential future scenarios in general–a critical feature nowadays.
For example, a historian might merge records from different archives to reconstruct an accurate historical narrative. This requires assessing the quality, validity, and reliability of various information streams, and ensuring consistency with existing knowledge. It’s also crucial to communicate transparently why certain information is deemed more reliable and why specific sources are considered credible. This process is akin to how financial businesses manage diverse data streams to generate more accurate and actionable insights.
By integrating information from multiple sources, financial organizations can create cohesive systems that enhance decision-making, improve efficiency, and maintain operational consistency. Manually applying logic and transforming this information into actionable knowledge would demand significant time and resources, which few organizations can afford. Automated technologies must therefore provide a robust framework that enables knowledge synergy through logical reasoning rather than relying solely on past statistics. This approach allows financial businesses to utilize their data more effectively, resulting in better strategic outcomes and a stronger competitive advantage.
Intelligent Technologies for Informed Decision-Making in Financial Businesses
Intelligent technologies–such as AI’s statistical methods and, in particular, Machine Learning (ML) algorithms–have demonstrated great efficiency in tasks where supervised learning (learning with labeled data) is applicable. This includes applications across a wide range of industries, offering significant advantages in fields such as language models and computer vision. However, challenges emerge when faced with high complexity and uncertainty; where data is often diverse, unstructured (or inhomogeneous in technical terms), and comes from sources with varying levels of reliability. This is typical for knowledge synergy tasks. Additionally, the “black-box” nature of many ML models can lead to inaccuracies [1] or biases [2] that are difficult to detect, especially when the data itself is flawed. In such situations, relying solely on traditional ML methods often proves insufficient.
In more complex environments, like fraud detection, where data may be incomplete or labels are missing (since accurately identifying fraud is challenging), supervised ML becomes less effective. To manage this, a “human-in-the-loop” process is frequently required. This means human experts are needed to assess data quality, label training data, and provide feedback, yet this approach can be both slow and expensive.
To overcome these limitations, more advanced systems combine ML with human oversight and probabilistic rule-based AI frameworks [3, 4]. These hybrid approaches extend beyond traditional methods by integrating supervised ML with rules-based reasoning (simplified as IF… THEN… rules, with probability), alongside other techniques (like inference from knowledge graphs). This allows to reason about uncertainty in environments with complex, structured relationships, which is often difficult using traditional machine learning techniques.
When data sources conflict or are incomplete, probabilistic rule-based AI systems excel by assigning importance, or “weights”, to different rules and allowing transparent fine-tuning (adjusting existing or adding new rules) clarifying which knowledge should be prioritized. This transparency helps organizations resolve data discrepancies more effectively, leading to better decision-making. In essence, these AI systems offer a blend of clarity and precision, providing a flexible framework for scalable and reliable decision-making [5].
Such interpretable, “white-box” approaches are particularly useful in tasks that require logical reasoning, like handling unlabeled data or simulation of future scenarios for forecasting. These advanced systems not only leverage historical data but also allow for real-time insights and expert input (if a case requires), enabling more informed and scalable decision-making.
Understanding Knowledge Synergies
The concept of knowledge synergy, also known as multi-source data fusion, involves integrating data from various sources to generate a comprehensive response to a specific request. This process entails collecting relevant information from diverse origins and combining it using appropriate methodologies to achieve a thorough and detailed understanding of a given topic. A graphical representation of a knowledge synergy engine that performs this task is illustrated in the figure below:
Knowledge synergy should not be confused with data mining, which often focuses on extracting valuable patterns from large datasets with appropriate methods and algorithms, effectively representing the results [6]. Data mining is also crucial for addressing various business needs, such as statistically predicting future values and understanding past results. Yet, it is a more statistical (rather than reasoning based on the acquired data) and less transparent process, serving a different purpose from a knowledge synergy such as pattern and value identification.
In this context, a highly credible and verified source serves as the ground truth, providing a reliable baseline for the information. Other sources, which may be less reliable, act as supplementary indicators that help to inform and enrich the established ground truth. By leveraging the strengths of both credible and less credible sources, multi-source data fusion enhances the overall quality and depth of the information, leading to more accurate and informed outcomes.
The concept of a knowledge engine [7] is not entirely new, that has been discussed in the literature for decades, but it is even more crucial today. The emphasis on source credibility and the reliability of data remains fundamental for effective, data-based multi-source decision-making. This aspect is particularly critical today, as challenges related to misinformation, disinformation, and intentional data corruption have become increasingly prevalent.
In this regard, interpretable and explainable systems are indispensable. They help clarify the factors influencing decisions, which is especially critical in fields that require legal compliance, such as anti-money laundering [8], fraud prevention [9], and credit-risk assessment [10]. These systems not only improve the reliability of outcomes but also ensure transparency, making the decision-making process clear and justifiable.
Other tools, such as Large Language Models (LLMs), can assist in identifying and gathering the necessary data. It is expected that, in the future, the power of LLM and artificial neural network (ANN) technologies will be harnessed for these and other operations, particularly when combined with probabilistic rule-based frameworks to enhance reasoning based on the underlying data. However, at present, these models are primarily suited for statistical analysis rather than producing detailed understanding through reasoning.
For financial institutions aiming to stay competitive, effectively managing and integrating diverse data sources is essential. Industry thought leaders highlight the importance of leveraging these resources to mitigate risks, enhance value, and improve product offerings. To meet these objectives, technology must be designed to be interpretable, adjustable, and adaptable to new contexts, ensuring it remains accurate and relevant over time.
Advantages and Impact of Synergies
Knowledge synergies provide organizations with substantial advantages by enabling more accurate insights and better-informed decision-making. The tech community often uses the phrase “data is air” to highlight how integral data is in our lives. We continually “inhale” data by analyzing incoming information and “exhale” it by generating data traces that can be analyzed further.
Practical Applications and Human-in-the-loop
By integrating information from multiple data sources, organizations can both strengthen their existing knowledge and uncover new, unexpected patterns. This comprehensive approach enhances understanding, aids in risk mitigation, and improves decision accuracy. Automating the integration process further streamlines decision-making by reducing manual reviews and addressing latency issues. For instance, in the financial sector, a company might use automated multi-source data fusion to detect fraudulent transactions in real-time. By analyzing data from transaction records, user behavior, and external fraud indicators, the system can quickly identify suspicious activities and trigger alerts, reducing the time needed for manual investigation.
Answering critical questions and fine-tuning of models still benefit from a human-in-the-loop approach. Experts can leverage their specialized knowledge to refine algorithms and improve outcomes. For example, in anti-money laundering (AML) systems, while automated processes can flag potentially suspicious transactions, financial experts review these flagged transactions to apply contextual knowledge and make final determinations. Probabilistic rule-based frameworks can play a key role here by being capable to combine probabilistic reasoning with explicit rules to enhance the system’s ability to handle complex and ambiguous situations, thus supporting more accurate and nuanced decision-making.
Leveraging Customer Data and Data Integration
Financial institutions often hold extensive information about their customers, including preferences and behaviors, which can significantly enhance various aspects of their operations. By effectively leveraging this data, organizations can improve areas such as tailoring product offerings and developing more precise risk mitigation strategies. One key application is in Know Your Customer (KYC) procedures, which involve verifying a prospective customer’s financial stability.
Going further, some institutions and researchers explore alternative data sources, such as social media networks, to gain a more comprehensive view of customers. Others point to its ethical and quality concerns, as the data from these sources can sometimes be unreliable or biased [8].
Integrating data from diverse sources, including alternative ones, presents challenges related to biases, accuracy, and data quality. Nonetheless, with appropriate mechanisms and processes in place, organizations can successfully navigate these issues. Effective data integration involves not just collecting and merging information but also applying robust reasoning and knowledge conflict resolution to draw actionable insights, automate processes, and scale operations.
Knowledge Synergy for Mitigation of Financial Risk
Mitigation of Financial Risk and Know Your Customer (KYC) compliance is essential for maintaining the security and integrity of financial transactions. More broadly, this concept can be viewed as assessing a customer’s financial credibility by understanding their financial behaviors and characteristics. The process not only benefits customers but also strengthens financial institutions by ensuring that they engage with legitimate and suitable clients. Effective financial credibility assessment of customers helps verify customer identities, assess their suitability for financial products, and understand their financial activities to prevent fraud and other illegal activities.
In the following, we discuss three critical aspects of KYC where knowledge synergy — integrating and utilizing diverse information sources — plays a pivotal role: account openings, financial credibility assessments for existing customers, and fraud mitigation.
Account Openings
For banks and large financial institutions, the KYC process during account openings is crucial. It involves verifying the identity and legitimacy of potential customers to ensure they are not involved in illicit activities such as money laundering or terrorism financing. The integration of multiple data sources, such as government databases, credit reports, and biometric data, enhances the accuracy of identity verification. This comprehensive approach reduces the risk of onboarding fraudulent or high-risk clients, thereby protecting the financial institution from potential legal and financial repercussions.
Financial Credibility Assessment
Financial credibility assessment for existing customers is vital for lending and loan institutions. It extends beyond evaluating the information provided by customers to a more in-depth analysis of their financial behavior and history. This involves integrating various data points, such as credit scores, transaction histories, and income statements, to form a complete picture of an individual’s or entity’s financial standing.
A significant, but not as commonly discussed, challenge in this area is managing the temporality of financial data. Financial institutions need to continuously update their understanding of a customer’s creditworthiness as circumstances change. This requires a dynamic, ongoing evaluation process rather than relying on a static, one-time assessment. Advanced analytics and real-time data integration are crucial for adapting to changes in a customer’s financial situation, ensuring that lending decisions are based on the most current and accurate information.
Let’s consider an example: when looking at a manufacturing company, day-to-day financial transactions might appear stable and normal. However, issues within their supply chain — such as delays or disruptions with suppliers — could signal potential financial trouble before these problems are evident in the company’s transaction data. By implementing a multi-source assessment system, a bank can gain early insights into these emerging issues and adjust their credit assessment accordingly, ensuring lending decisions are based on the most up-to-date and comprehensive information. Such an application is similar to what a human analyst would do, equipped with large amounts of data.
Fraud Mitigation
Fraud prevention is a critical concern across all types of financial institutions, from banks to investment firms. Effective KYC processes can significantly enhance fraud detection and prevention efforts. By leveraging a multitude of data sources — including transaction records, behavioral patterns, and external fraud indicators — financial institutions can better identify suspicious activities and potential threats.
Analyzing patterns in transaction data alongside customer behavior and external fraud signals enables institutions to detect anomalies that might indicate fraudulent activities. This multi-faceted approach improves the accuracy of fraud detection systems, reduces the number of false positives, and enhances overall security.
A documented case in the Italian Financial Intelligence Unit (FIU) describes a fraudulent activity involving a network of relatives engaged in money laundering [8]. Using a hybrid system of ML and knowledge graph reasoning with rule-based logic, the FIU identified suspicious behavior by analyzing ownership patterns and members’ relationships, revealing a family circle that collectively controlled a bank through a complex shareholding structure. This system uncovered how the family members, by obscuring their ultimate ownership, attempted to launder illicit funds through a fake loan issued by a bank they controlled.
Conclusion
In conclusion, knowledge synergy offers substantial advantages in financial decision-making processes across various domains. By integrating data from multiple sources and applying flexible rule-based reasoning, organizations can achieve a more comprehensive and accurate understanding of complex situations. In such environments, hybrid approaches that combine machine learning, human oversight, and probabilistic rule-based frameworks are essential. These approaches effectively manage unstructured or conflicting data to address challenges in fraud detection, financial credibility assessments, and mitigation of financial risk, leading to more informed and precise outcomes. This is particularly valuable in finance, where accurate decision-making is critical, but the potential applications extend to other sectors like self-driving cars, insurance, and medicine [11], where real-time analysis and robust decision-making are essential.
As financial institutions increasingly seek efficient intelligent technologies, these hybrid systems help manage risks, offer tailored solutions, and ensure compliance, all while maintaining operational consistency and securing long-term success in an evolving financial landscape.
Bibliography:
1. Liang, W., Tadesse, G.A., Ho, D. et al. Advances, challenges and opportunities in creating data for trustworthy AI. Nat Mach Intell 4, 669–677 (2022). https://doi.org/10.1038/s42256-022-00516-1
2. Bowen III, D. E., Price, S. M., Stein, L. C. D., & Yang, K. (2024). Measuring and mitigating racial bias in large language model mortgage underwriting. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4812158
3. Koller, D., & Friedman, N. (2012). Probabilistic graphical models principles and techniques Daphne Koller; Nir Friedman. MIT Press.
4. Domingos, P., & Richardson, M. (2007). Markov Logic: A Unifying Framework for Statistical Relational Learning. In Introduction to Statistical Relational Learning (pp. 339–371), eds. Getoor, L., Taskar, B. MIT Press. Cambridge Mass.
5. Schaefer, T. (2023). In need for both accuracy and interpretability? give probabilistic rules a try. Medium. https://medium.com/stratyfyai/in-need-for-both-accuracy-and-interpretability-give-probabilistic-rules-a-try-6713a74c7776
6. Gullo, F. (2015). From patterns in data to knowledge discovery: What data mining can do. Physics Procedia, 62, 18–22. https://doi.org/10.1016/j.phpro.2015.02.005
7. Yager, R. Some Considerations in Multi-Source Data Fusion. In: Ruan, D., Chen, G., E. Kerre, E., Wets, G. (eds) Intelligent Data Mining. Studies in Computational Intelligence, vol 5. Springer, Berlin, Heidelberg. https://doi.org/10.1007/11004011_1
8. Bellomarini, L., Laurenza, E., & Sallinger, E. (2020). Rule-based Anti-Money Laundering in Financial Intelligence Units: Experience and Vision. In Proceedings of the 14th International Rule Challenge, 4th Doctoral Consortium, and 6th Industry Track @ RuleML+RR 2020 co-located with 16th Reasoning Web Summer School {(RW} 2020) 12th DecisionCAMP 2020 as part of Declarative {AI} 2020, Oslo, Norway (virtual due to Covid-19 pandemic), 29 June — 1 July, 2020 (pp. 133–144). http://hdl.handle.net/20.500.12708/58323
9. Islam, S., Haque, Md. M., & Rezaul Karim, A. N. (2024). A rule-based machine learning model for Financial Fraud Detection. International Journal of Electrical and Computer Engineering (IJECE), 14(1), 759. https://doi.org/10.11591/ijece.v14i1.pp759-771
10. Tian, M., Li, H., Huang, J., Liang, J., Bu, W., & Chen, B. (2022). Credit risk models using rule-based methods and machine-learning algorithms. Proceedings of the 2022 6th International Conference on Computer Science and Artificial Intelligence. https://doi.org/10.1145/3577530.3577588; Rudin, C., & Shaposhnik, Y. (2019). Globally-consistent rule-based summary-explanations for machine learning models: Application to credit-risk evaluation. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.3395422
11. D’Haese P-F, et al. (2021) Prediction of viral symptoms using wearable technology and artificial intelligence: A pilot study in healthcare workers. PLoS ONE 16(10): e0257997. https://doi.org/10.1371/journal.pone.0257997
Interested in more original insights from our team? Subscribe to our quarterly newsletter.