Week 5: From a Phishing Prediction to a Security Decision

This week in CYBR 325, I trained a Random Forest classifier on the PhiUSIIL phishing dataset in Google Colab. The displayed metrics rounded to 1.0000, but the confusion matrix showed one missed phishing website among 47,159 test examples. Accuracy was approximately 99.998 percent. For our Security Operations topic, that result raises a practical question: what should an organization do when a detector misses an attack?

I selected Phishing Guidance: Stopping the Attack Cycle at Phase One, released by CISA, NSA, the FBI, and MS-ISAC in October 2023. The guide distinguishes phishing intended to steal login credentials from phishing used to deliver malware. It addresses prevention and incident response, with protections including phishing-resistant multifactor authentication, filtering malicious links and attachments, and protective DNS (Cybersecurity and Infrastructure Security Agency et al., 2023). Although it predates this course, its layered approach provides a useful way to examine where my classifier could fit.

My connection is that a phishing prediction would be one input to a security decision. A false positive could block a legitimate website and interrupt someone’s work. A false negative could let a dangerous page reach a user. My test produced no false positives and one false negative, but those counts describe this test set. They do not establish how the model would perform against future attacks. Other protections would still matter if its classification were wrong.

The experiment also placed limits on what I could claim about feature engineering. I created three features representing relationships among URL characteristics, but I did not train a baseline model without them. None appeared among the ten most important features in this run. I therefore cannot attribute the high score to my additions. Before using a model operationally, I would want evidence that a change improves detection, along with testing on data that better represents its intended environment.

Our separate Group 1 lab explored deepfake detection and how prompt engineering can produce more specific cybersecurity information. I served as presenter and timekeeper. That work adds another question to my thinking about security operations: how should a team verify an AI-assisted finding before acting on it? A detector’s output needs a review and response process, whether it concerns a suspicious website or manipulated media.

In my assignment’s ethics analysis, I proposed limiting sensitive telemetry, restricting access, monitoring errors, and providing human review for consequential blocking decisions. These responsibilities would continue after deployment. My main takeaway is that building a classifier is only part of building a useful security capability. Its value depends on how people interpret its output, respond to mistakes, and maintain protection as threats change.

References

Cybersecurity and Infrastructure Security Agency, National Security Agency, Federal Bureau of Investigation, & Multi-State Information Sharing and Analysis Center. (2023, October 18). Phishing guidance: Stopping the attack cycle at phase one. https://media.defense.gov/2023/Oct/18/2003322402/-1/-1/0/CSI-PHISHING-GUIDANCE.PDF

AI use disclosure

ChatGPT was used at Level 3, AI Collaboration, to assist with source discovery, organization, drafting, and revision. The lab activities, model results, and course connections in this draft are based on my submitted assignment and lab summary.

Leave a comment