Sunday, February 9, 2025

From Innovation to Irrelevance: Reflections on the Paper "The Intelligence Curse"

Last month an article was published in Less Wrong titled "The Intelligence Curse" by Luke Drago. Luke Drago who works on AI Governance projects has written this piece that everyone should read about a possible future that is quite depressing in its implications - that AI, perhaps humanity’s greatest innovation could make most of humanity irrelevant. But although the ideas in the paper do have a non-zero chance of occurring, it's one that I don't think is likely and in this post I'll explain why.

I almost titled this post: "Why the Future Will Still Need Us" as a nod to the seminal article written in 2000 from Wired titled "Why the Future Doesn't Need Us." In that article, Bill Joy, co-founder of Sun Microsystems, expressed deep concern over the potential dangers posed by emerging technologies; specifically genetics, nanotechnology, and robotics (GNR). He argued that these advancements could surpass human control, leading to self-replicating entities capable of causing massive destruction. Interestingly, in that 2000 article, AI wasn't at the top of his existential threats like it is now for the P(doom) crowd.

Other important articles that are more recent and AI specific that I think are very relevant for this discussion that I commented about here and here on are "My Last Five Years of Work" by Avital Balwit and "Situational Awareness" by Leopold Aschenbrenner. Both of these authors address the societal transformations with the advent of artificial general intelligence (AGI). Balwit writes on the potential obsolescence of human labor, suggesting a future where AGI could render traditional employment obsolete. Similarly, Aschenbrenner discusses the rapid progression towards AGI, emphasizing the need for heightened awareness and preparedness for the imminent changes. Both emphasize the urgency of addressing the societal and ethical implications of AGI’s integration into various facets of life. Likewise, this new paper by Luke Drago is a continuation of this thread of a need to think about and prepare for a possible future that could be radically different from today.

But before I get into the "The Intelligence Curse" paper itself, I want to state that I am an AI optimist and believe in accelerating AI progress and that AI acceleration can be done in a way that is beneficial and responsible. Accelerating change in and of itself can be disorienting. But I believe AI may give us the only real hope to counter and control forces in the 21st century that have truly existential implications such as global warming, conflicts over energy and resource scarcity, food scarcity and malnutrition, disease, pandemics, etc.

So what is the "Intelligence Curse"?

The article introduces a scenario where AGI fundamentally alters economic incentives, leading to widespread human irrelevance and reduced investment in human welfare. Unlike past technological revolutions that expanded human potential, AGI is more like a concentrated natural resource such as oil. Its control will rest with a few powerful actors, corporations and governments, who prioritize maximizing returns from AI systems over investing in people.

In this post-AGI economy, companies would replace human labor with AGI because these AI agents are faster, cheaper, and more reliable. This shift removes the economic motivation to invest in traditional areas like education, infrastructure, and social welfare, mirroring the “resource curse” seen in "rentier" states. For example, the Democratic Republic of Congo (DRC) is rich in natural resources, with an estimated $24 trillion worth of untapped minerals, yet most of its population lives in poverty. Similarly, the "Intelligence Curse" could cause states and corporations to focus solely on extracting value from AGI, leaving human development and welfare behind.

The social and economic consequences of this transformation could be dire. As AGI replaces human labor across all fields, most people would lose economic relevance. States may increasingly rely on taxing corporations instead of individuals, while AI labs become powerful rent seekers, controlling significant economic resources. Consequently, social safety nets and public funding for human-centered initiatives would likely diminish, leaving many economically vulnerable.

The author emphasizes that incentives drive societal decisions. Today, states and corporations invest in human capital because they derive economic returns from people’s productivity. Education, infrastructure, and social programs create skilled workers, productive citizens, and consumers who fuel the economy. This feedback loop benefits those in power. However, once AGI becomes the primary driver of economic value, this incentive disappears. Powerful actors will no longer need human labor or consumer demand, relying instead on scalable, cost-effective AI systems.

He argues that this shift will lead to widespread disinvestment in people, collapsing the economic structures that support education, social programs, and infrastructure. Without an economic reason to invest in human capacity, powerful actors will focus on optimizing and extracting value from AI, creating a society where the majority of people become irrelevant to the system’s success.

So before, I get to why I think this scenario is unlikely, I'll extend out his dystopic ideas even more. Even looking at this future, there would be some optimists who would say that we would be living in a post scarcity society and that everything would be freely available and what is not freely available could be obtained by a government implementing Universal Basic Income (UBI). It will be utopia! Even if all humanity is beholden to authoritarian governments and corporations. But in Luke Drago's argument that centers solely around incentives, what would be the incentive to provide UBI? Altruism?

Governments in this scenario would only provide enough income to ensure that people don't riot and that those in power can keep that power. AI enabled feudalism. Those corporations and governments in control will be in control of AI that will be more persuasive than the best advertising campaign or the greatest politician that's ever existed. So in this post scarcity society, with endless AI provided entertainment, AI persuasiveness, and advancements in pharmaceuticals that can enable whatever human beings want to experience.

But what about future employment in this dystopia? Human beings derive meaning from doing "meaningful" things. We can turn to the article by Avital Balwit of "My Last Five Years of Work" I mentioned above and in my blog post about that article to see that humans would still find meaningful things to do. But would these corporations and governments still need to hire people to do specific jobs? There would still be those who work directly for these large corporations and governments - although these positions would grow fewer and fewer as AI is given more control over these tasks. What about those professions that involve human to human experiences: nurses, caretakers, physical therapists, etc. Yes, there will continue to be a need for those professions. Also, human created entertainment that can be experienced like plays, athletics, musical performances would be popular. Human created art and literature that could only be based on real, actual human experiences would be popular. So books based on human experiences that are autobiographical or semi-autobiographical would also be popular (think Dave Eggers book A Heartbreaking Work of Staggering Genius). So sure, there will still be these types of human activities, but they would be a small percentage compared to that of current employment numbers.

And to put a really cynical point on this dystopic vision, the greatest employer would be military services and security forces. Governments would still need large militaries in their competition with other AI equipped governments. And both governments and corporations would need large security forces for protection and pacification purposes from the broader population.

Okay, okay this has all been very depressing, because the author makes the case that this future is an inevitable outcome of AI advancement unless specific policies and actions are taken - and to be fair he does promise a follow up article to outline more on those specifics. But I want to outline why I think this future is not inevitable and is actually unlikely.

1) Timeline and Transition Period

First let's look at his timeline and the nature of that timeline. I'm going to quote two paragraphs in full here, so we have his full context:
I believe that artificial general intelligence (AGI), specifically “a highly autonomous system that outperforms humans at most economically valuable work” is technologically achievable and >90% likely to exist in the next 1-20 years (and honestly, 10 years feels way too long). You should too.

Once AI systems that are better, cheaper, faster, and more reliable than humans at most economic activity are widely available, the intelligence curse should begin to take effect. We should expect to be locked into the outcome 1-5 years after this moment.

Let's examine this timeline. Although the definition of AGI is rather nebulous depending on who you talk to, we can look at the most recent predictions of some of the leaders in the field. Sam Altman of OpenAI now says it could be this year. Dario Amodei the CEO of Anthropic thinks that 2026/2027 AGI could happen. People like Demis Hassabis of DeepMind and Yann Lecun of Meta have a longer timeline with Yann Lecun generally disliking the term AGI in favor of human-level intelligence. But it's fair to say that regardless of the definition of the timeline, predictions have been decreasing.

I generally think the shorter timelines are correct of 2025-2026, but let's say the timeline when most reasonable people can agree that AGI has arrived is 5 years from now. Then according to Luke Drago, this dystopic scenario would be "locked in" within 1-5 years after that. Let's take the median of his estimate of 3 years, so then in 2030 AGI widely exists and in 2033 we have this dystopia of the intelligence curse. But if it is much sooner - if Sam Altman and Dario Amodei predictions of AGI are correct, then this dystopia arrives before 2030.

Let's look at the World Economic Forum report just released in January of 2025 and I'm going to quote directly from the report as to how it was compiled and the timeframe for its forecast:
The Future of Jobs Report 2025 brings together the perspective of over 1,000 leading global employers—collectively representing more than 14 million workers across 22 industry clusters and 55 economies from around the world—to examine how these macrotrends impact jobs and skills, and the workforce transformation strategies employers plan to embark on in response, across the 2025 to 2030 timeframe.

It is looking at the timeframe we are interested in:

Extrapolating from the predictions shared by Future of Jobs Survey respondents, on current trends over the 2025 to 2030 period job creation and destruction due to structural labour-market transformation will amount to 22% of today’s total jobs. This is expected to entail the creation of new jobs equivalent to 14% of today’s total employment, amounting to 170 million jobs. However, this growth is expected to be offset by the displacement of the equivalent of 8% (or 92 million) of current jobs, resulting in net growth of 7% of total employment, or 78 million jobs.

So over this period when AGI could very well be happening and the intelligence curse should be wrecking havoc with economies, the World Economic Forum is predicting a net growth in jobs. To be sure, it lists out several major transformative changes that will be taking place even beyond AI - demographic changes (population aging), climate change, "geoeconomic fragmentation and geopolitical tensions", but even with these changes there will still be positive increases in employment. And I don't think this increase in employment is fully taking into account opportunities that will be created by AGI in technologies that cannot be possibly imagined at this point.

This hardly sounds like the dire employment scenario described in the paper.

There is also the problem I have with this paper and also with the "My Last Five Years of Work" paper is that it is strongly implied that once AGI is achieved that there is a very quick transformation - a complete makeover of economic and social structures. But there are all kinds of inertial forces at work that prevent radical transformation across industries to happen quickly. These difficulties include changing highly dependent systems, bureaucracies, complicated logistics, regulation conformities, and just the natural resistance to change from people across organizations. Often the large costs and timelines of changing complicated systems will drag out AGI implementation in parts of the economy such that AGI will be very unevenly distributed.

There will be this transition period after AGI is achieved where some automations will be easy and quick to do, but other parts of the economy could take over 20 years. And it's because of this transition period - and if governments and organizations use this transition period well, then that time to plan could prevent the dystopia of the intelligence curse.

2) Paradox of AI Alignment

My idea of the paradox of AI alignment is that we are striving for an AI that is aligned with human values and we spend a lot of time and effort trying to create that alignment, but in the end we may as a species be much better served if AI alignment proves to be impossible.

Over the last couple of years there have been highly publicized departures of very notable alignment people leaving OpenAI and recently at Anthropic to go elsewhere to do alignment research. The stories around these departures is ostensibly that the companies that they were at weren't serious enough about AI dangers, were moving too fast, and weren't giving alignment teams enough resources. But I think the reality is that AI alignment is just really, really hard and the frustration is that it has been very difficult to make progress and sustain progress in alignment - especially when AI is constantly changing. AI alignment, in the end, might just be impossible. However, alignment being impossible could be a good thing.

Here's why:

In a world with an "aligned" AGI/ASI we would be trusting it more, giving it more general objectives, objectives that would be more fuzzy, where the AI could choose from more and more paths to achieve those objectives. We would be letting it be more autonomous to achieve those general objectives, because it's "aligned" with human values. In a world, where we trust AI, there's less need for human oversight of AI, human management of AI, human security of AI, human deployments of AI. But without trusted alignment, AIs will need human partners, AI will need to be directed to specific tasks and with that need for management that ensures that a whole sector of jobs will continue to grow if we assume that we would never have AI fully aligned.

Now I know that alignment won't be an all or nothing proposition, that we can have alignment that does some goals well, but as long as we don't have fully trusted AI, I believe we will be better off. And since alignment is difficult - that's a good thing.

3) Myth of AI Centralization

By the myth of AI centralization, I am referring to a widely held belief by almost everyone, until maybe very recently, that in order to create and control AI you needed enormous resources to train and deploy these models. Estimates on training GPT 3.0 were around $4-5 million, however the estimated cost to train a model the size of GPT 4.0 is around $63 million. The data centers have gotten larger, the number of GPUs needed have increased, and the energy requirements have increased to the point that major labs talk about having their own nuclear power plants. And just recently, it was announced that Stargate would be a data center costing $500 million that would begin construction in the middle of 2025.

And Luke Drago makes this his main point for comparing AGI to resouces like oil in rentier states. He states:

But AGI looks a lot more like coal or oil than the plow, steam engine, or computer. Like those resources:

  • It will require immensely wealthy actors to discover and harness.
  • Control will be concentrated in the hands of a few players, mainly the labs that produce it and the states where they reside.

But almost at the same time as Stargate was being announced with its enormous price tag, Deepseek R1 was released and open sourced, rivaling the best models at the time at a fraction of the size and cost. Regardless of if the cost was not exactly what was first reported, it is still much smaller than what the large labs have been saying they needed. The combination of ideas like distillation, improved data sets, improved reinforcement learning techniques, open source, and even potentially different architectures than transformers are leading to the possibilities of smaller models. Plus, chips should continue to drop in price to the point that I believe very small groups - and individuals will be able to create their own AGI versions.

And this I believe is the biggest wild card that is not being accounted for in the "The Intelligence Curse" paper. It will not be a few mega-trillion dollar companies controlling AGI. Instead AGI will be comoditisized and democratized. Individuals will be empowered to control intelligence in ways that they want and not as dictated by large AI labs.

Conclusion

While The Intelligence Curse presents a sobering and dystopian vision of a post-AGI future, I am unconvinced that this outcome is inevitable or even likely. The future is rarely as linear as such scenarios suggest, and history shows us that technological transitions are complex, uneven, and full of surprising turns. The timeline proposed by the paper feels overly compressed, underestimating the inertia of existing systems and the time it takes for society to adapt. Moreover, the assumption that AGI will remain in the hands of a few centralized actors ignores recent developments in open source AI and more efficient models that could democratize access to powerful technologies.

The paradox of AI alignment may also work in humanity's favor, ensuring that humans remain an integral part of managing and guiding AI systems for the foreseeable future. Far from eliminating the need for human oversight, imperfect alignment could create new categories of work that preserve human relevance in a world increasingly shaped by AI. Additionally, the myth of AI centralization overlooks the potential for decentralized, individual driven innovation that could disrupt the monopolistic control envisioned in the paper.

Ultimately, while it's crucial to think critically about the risks posed by AGI and advocate for responsible AI development, it’s equally important to remain grounded in what we know about technology adoption and economic change. The future will still need us - not just as passive recipients of technological progress but as active participants in shaping its direction.

I'm convinced there's reason to believe that AI will enhance human potential rather than diminish it.

Sunday, January 12, 2025

Elements of Monte Carlo Tree Search - Typical and Non-typical Applications

Monte Carlo Tree Search (MCTS) offers a very intuitive way of tackling challenging decision making problems. In essence, MCTS combines the structure of a tree search with the power of random sampling to navigate large, complex search spaces effectively. I've been interested in MCTS for a long time - this interest was cemented in 2016 by the victory of Alpha Go over Lee Sedol and the famous "Move 37." By the way, I wrote an entire post about Move 37 and the role of innovation in AI here. Move 37, a move that was thought at best to be "unique" but was at the moment the move was played thought to be a horrible move - a move that no respectable Go player would make. But the move was both effective and innovative. It was this balance between exploiting an immediated advantage and exploring the space for potentially even better solutions that gave Alpha Go the victory in a game that because of its complexity and need for imagination was thought impossible for AI. This exploitation vs exploration dynamic is the crux of MCTS and it along with neural network evaluations and reinforcement learning gave Alpha Go the victory.

In this post, I want to talk about the main ideas behind MCTS, where MCTS can be used, and then walk through the code for implementing a MCTS in a game of Connect 4. But if you want to skip right to playing the game: it is here and the Github code is here.

So what kind of applications is MCTS appropriate for? Historically the typical scenario is around games and specifically finite two-person zero-sum sequential games. In other words, in these types of games, there are two players competing directly against each other, the gain of one player is exactly the loss of the other (zero-sum), the game progresses through sequential moves, and the game has a finite number of states and actions, leading to a clear end.

These two player games are the typical application of MCTS, but I've long been interested in applications beyond these types of games. Any problem where there is a large search space, that has tree like decision nodes that operate sequentially under limited resource or time constraints. Any kind of optimization problem with these characteristics MCTS could be applicable. Robotics is an obvious example. A robot has to complete some objective, but there are many different paths to complete that objective. But how about some not so obvious applications. In education, MCTS could help in what I'm calling Personalized Learning Path Optimization (PLPO). A PLPO could as a student engages with learning activities customize a personalized curriculum for a student in an adaptive learning platform, balancing engagement, challenge, and skill acquisition. In this situation, the state represents the student’s current skill level, engagement, and performance history. Actions include presenting different types of learning activities (e.g., videos, quizzes, projects). The MCTS simulations predict the student’s response to a sequence of learning activities, optimizing for long-term improvement and retention.

In biotech, an example would be to optimize the sequence of biological reactions in a synthetic pathway for producing potential drugs. The states would represent the current pathway, including enzymes and intermediates. Actions involve adding, removing, or modifying reactions in the pathway. Then MCTS simulations predict yield, efficiency, and stability, optimizing for production goals and feasibility.

Another non-typical example I want to talk at length about is optimizing a customer journey designed for e-commerce platforms. In this example, the goal would be to guide consumers through a personalized journey (e.g., product recommendations, promotional offers, and website interactions) to maximize conversions, repeat purchases, and customer satisfaction.

I want to come back to this example of using MCTS in optimizing the customer journey, but first I want to go into a larger explanation of how MCTS works in the typical case of two person games.

MCTS combines the structure of a tree search with random sampling to navigate large, complex search spaces. Unlike traditional search methods that often rely on complete or near complete look-ahead, MCTS focuses on sampling potential outcomes and incrementally refining its estimates of how good or bad certain moves or actions are. This makes MCTS particularly appealing in domains where exhaustive searches become prohibitively large.

One way to understand MCTS is by tracing the iterative loop through four stages:

  1. Selection
  2. Expansion
  3. Simulation
  4. Backpropagation

Each of these stages contributes to building up a more accurate picture of the search space. During selection, the algorithm starts at the root of the search tree (representing the current state) and travels down along existing paths, choosing actions that optimize the balance between exploitation (picking actions that have worked well so far) and exploration (investigating actions that remain relatively untested). A common way to manage this trade off is by using the Upper Confidence Bound applied to Trees (UCT) formula. If \( N \) is the total number of visits to the current node and \( n \) is the number of visits to a child node, while \( \overline{X} \) is the current estimate of the child’s average reward, then the child node’s UCT value can be computed as:

\( \text{UCT} = \overline{X} + c \sqrt{\frac{\ln(N)}{n}} \)

\( \overline{X} \) = current estimate of the child’s average reward
\( c \) = constant controlling the exploration-exploitation balance
\( N \) = total number of visits to the current node
\( n \) = number of visits to a child node

Here a larger \( c \) encourages more exploration of unexplored moves. The goal of this formula is to favor actions that have high average rewards while still allocating some attempts to actions with few visits (because they might lead to surprising or better results once explored more thoroughly). This is how we get moves that might look to a bystander as "unusual" or "non-standard."

After reaching a leaf node, either a node that has not been explored yet, or one that is not fully expanded with potential child nodes, the algorithm proceeds to the expansion phase. In this stage, it creates one or more child nodes corresponding to unvisited actions. This expansion step gradually grows the search tree over time, ensuring that new parts of the state space are discovered rather than staying confined to what has already been visited.

The third stage, simulation, or you will often hear it referred to as the playout, is where the “Monte Carlo” part comes into play. From one of the newly added child nodes, MCTS simulates a sequence of moves (often random or guided by a light heuristic) all the way to a terminal outcome. For example, in a game, that outcome might be a win, a loss, or a draw. In a more general planning or scheduling context, it could be a success or failure, or a particular reward value. These random playouts, repeated many times, serve as samples of what might realistically happen if a particular action path is chosen.

Once the simulation is complete, the algorithm performs backpropagation to propagate the result of the simulation back up the tree to the root. Along the path taken in the search tree, each node’s statistics - like how many simulations passed through that node, how many ended in a favorable outcome, or the average return are updated. Over many iterations, these accumulated statistics become increasingly meaningful, revealing which branches of the search tree appear promising and which ones do not.

MCTS is successful in many applications because it doesn’t require a specially devised evaluation function for intermediate states. Given enough playouts, the sampling can approximate the true value of each action. This versatility is why MCTS has been popular in board games like Go, where it formed the backbone of AlphaGo’s decision making system. It is also employed in video game AI, robotics for path planning, and even scheduling problems; anywhere the decision space is complex and uncertain.

Yet, MCTS is not without challenges. Its effectiveness depends significantly on the quality of the playouts - because remember these playouts are taken place under a constraint, which is often a time limit. If the simulations are purely random in a very large or complex domain, they might produce misleading estimates. For this reason, many implementations add a small amount of domain knowledge or heuristics to steer simulations toward more likely or more relevant outcomes (I will do this in the Connect 4 example below). Another concern is the size of the search tree itself: while MCTS is more scalable than exhaustive methods, it can still blow up in expansive problem spaces if it does not prune or limit growth effectively. Implementation details such as how states are represented, how child nodes are stored, and how results are backpropagated can significantly impact its performance.

In practice, the so-called “anytime” property of MCTS is one of its greatest strengths. You can let the algorithm run for as many or as few iterations as time allows, and it will offer the best solution it has found up to that point. Longer run times translate into more simulations, deeper exploration, and more refined estimates of action quality, but even with limited time, MCTS can generate a decent approximation of the optimal move.

So by blending random sampling with incremental, iterative updates to a tree of possibilities, MCTS circumvents the need for elaborate heuristics or full look-ahead. Through repeated applications of selection, expansion, simulation, and backpropagation, MCTS transforms what might otherwise be an intractable search into a manageable process, making it a "go to" strategy for complex game spaces like Go, Chess, and Othelo.

Before moving on to the Connect 4 code, I should mention the algorithm of minimax search with alpha/beta pruning. Historically, minimax has been the "go to" algorithm for games like Chess and the best Chess engines in the world have been built around minimax like Stockfish. Although recently a MCTS type Chess engine named Leela has competed very well with Stockfish (although technically it's not a pure MCTS - it uses a neural network to simulate the rollouts). Minimax with alpha/beta pruning is a very deterministic search. It systematically explores the game tree to a certain depth (or until terminal states). Typically, games like Chess with large branching factors must limit search depth. At the leaf nodes (or at the cutoff depth), an evaluation function (heuristic) estimates the position’s value, which is where pruning comes in. Pruning uses α (alpha) and β (beta) bounds to prune branches that cannot affect the final minimax decision.

The main difference between MCTS and minimax is that MCTS is probablistic and minimax is deterministic. Minimax is effective in games with reasonable branching factor like Chess and Checkers. But with games like Go, the number of states grows exponentially with depth of search.

So based on this discussion, you're probably thinking that minimax would be the best algorithm for Connect 4, since search space of possible moves and the depth of search is very limited - and you would be right. And certainly anyone building a game of Connect 4 would almost automatically use a minimax type algorithm. But with the following code using MCTS, I'm showing MCTS can do as well as minimax given a time constraint.

The Github repository can be found here. The most important function for the implementation is in the mcts.py file in a function unsurprisingly called mcts. This function steps through the four stages I outlined above of selection, expansion, simulation, and backpropagation. It iteratively does this under a time constraint. I have set my time constraint for 1 second - it has to come up with a move in under 1 second. In the selection stage, it is looking for the child node with the best UCT value. Once that node is found, the expansion stage expands that node by taking one untried move and creating a child node. Then the simulation stage simulates a random game of moves (rollout) from that child node until there is a winner or a draw. The backpropagation stage collects the simulation results up the nodes in that tree path. Once the while loop is finished, the node with the highest node count is the move that will be taken (which in the case of Connect 4 is the column that the piece will be dropped).

def mcts(root_board, current_player, simulations=500, time_limit=1.0):
    """
    Perform MCTS from the root state and return the column of the best move.
    """
    start_time = time.time()
    root_node = MCTSNode(root_board, current_player)
    
    while (time.time() - start_time) < time_limit:
        # 1. Selection
        node = root_node
        while not node.untried_moves and node.children:
            node = best_child(node)
        
        # 2. Expansion
        if node.untried_moves:
            node = expand_node(node)
        
        # 3. Simulation
        winner = simulate_game(node.board, node.current_player)
        
        # 4. Backpropagation
        backpropagate(node, winner)
    
    # After time is up, pick the child with the highest visit count.
    best_move_node = max(root_node.children, key=lambda c: c.visits) if root_node.children else None
    
    if best_move_node is None:
        # fallback if somehow no children
        return random.choice(get_valid_moves(root_board))
    
    # Return the column that leads to best_move_node
    for col in get_valid_moves(root_board):
        candidate_board = make_move(root_board, col, current_player)
        if candidate_board == best_move_node.board:
            return col
    
    return random.choice(get_valid_moves(root_board))

When the computer is looking to make its move, it doesn't automatically call this mcts function to get its next move. As I mentioned above, many games will include heuristics to steer the decision to look for a forced Connect 4 win or to block the opponent's forced win. Chess engines will incorporate special case checks for forced checkmates or forced draws and then do their algorithm search like the minimax. Similarly, I created a find_immediate_win_or_blockade function so as to not miss a forced win or loss that it will look at before doing the mcts function.

So this function will look to make sure the computer doesn't miss a move before it does its time budgeted move search:
def find_immediate_win_or_block(board, current_player):
    """
    Check if current_player can immediately win,
    or if the opponent can immediately win next turn (then block).
    """
    # Immediate win
    for col in get_valid_moves(board):
        temp_board = make_move(board, col, current_player)
        if check_winner(temp_board, current_player):
            return col
    
    # Block opponent
    opponent = get_next_player(current_player)
    for col in get_valid_moves(board):
        temp_board = make_move(board, col, opponent)
        if check_winner(temp_board, opponent):
            return col
    
    return None  
One more function I think I need to explain looked like the code below. This ai_move function would call the find_immediate_win_or_block function to look for an immediate win or block and if there wasn't an immediate win or block would then call the mcts function to look for a column to drop its piece. However, when I gave the link to family and friends, they complained that they couldn't ever win. So I added a 1 through 5 difficulty setting that the user can set. And if you look at the updated Github code, level 5 will do this original code, but levels 1-4 will choose a probabilistic amount of times that it will not do the MCTS and instead it will just drop a piece randomly. There are other ways to handle difficulty level, such as lessening the search time or space, but for Connect 4, mixing in a random move seems to work well.
def ai_move(board, current_player, difficulty):
    """
    Decide AI move:
    1. Check immediate win/block
    2. Otherwise use MCTS
    """
        
    col = find_immediate_win_or_block(board, current_player)
                  
    if col is not None:
        return col
                          
    return mcts(board, current_player, simulations=300, time_limit=1.0)
Initially, I wrote this Connect 4 code where the output was displayed in the terminal, but I wanted to share it with other people. But at the same time I wanted a minimalist implementation, so I did the unusual idea of doing it as a Streamlit application using emoticons. The end result is here, which you can play.


Non-typical MCTS: Customer Journey

Now that we have an explanation and example code of the typical application of MCTS two player game, we can talk about an example of non-typical application. There are many, many possible applications, but let's just look at one, which is an e-commerce platform where the goal is to guide consumers through a personalized journey (e.g., product recommendations, promotional offers, and website interactions) to maximize conversions, repeat purchases, and customer satisfaction.

In this type of example we have:
  • States in the customer journey:
    • Customer demographics and preferences.
    • Interaction history (e.g., clicks, time spent, abandoned carts).
    • Contextual data (e.g., time of day, device type, or browsing session).
  • Actions
    • Recommending specific products.
    • Sending a targeted promotional email or notification.
    • Offering discounts or free shipping.
    • Displaying upsell or cross-sell opportunities.
  • Reward Function
    • Conversion events (e.g., purchases or sign-ups).
    • Total revenue or profit generated.
    • Customer satisfaction and engagement metrics (e.g., session length, feedback).
  • Simulations
    • Rollouts simulate different sequences of actions to predict customer responses, using probabilistic models of consumer behavior (e.g., likelihood of purchase or churn)
  • Balancing of Exploration and Exploitation
    • MCTS balances exploring new strategies (e.g., trying novel product bundles or marketing messages) with exploiting known successful approaches (e.g., strategies that previously led to high-value purchases).

For example, a customer visits an online store and views a product but doesn’t add it to their cart. The platform needs to decide: should it recommend a similar product?, offer a discount on the viewed product? or highlight customer reviews or provide an FAQ link to address hesitation? Using MCTS, the platform simulates potential outcomes of these actions (e.g., conversion likelihood) and dynamically selects the one with the highest long-term reward, considering both immediate revenue and customer loyalty.

Using MCTS for the customer journey provides dynamic personalization, long term optimization, and is scalable and probabilistic. The MCTS driven system adapts to real-time customer interactions, creating highly tailored journeys. Instead of focusing on one off conversions, the system optimizes for lifetime customer value. MCTS handles the uncertainty and variability in consumer behavior effectively, especially for large scale platforms. Using a MCTS algorithm could enable smarter decision making and better shopping experiences.

Conclusion

Monte Carlo Tree Search (MCTS) is a robust and adaptable algorithm capable of addressing a diverse range of decision making situations. It has been successfully applied in domains such as board games, but could be applied atypically in areas like robotics, education, biotechnology, and e-commerce. This is primarily because of its versatility and efficacy in complex and uncertain search spaces. By balancing exploration and exploitation, MCTS optimizes decision making in contexts where exhaustive searches are computationally infeasible.

Additionally, there are many other variations of MCTS besides the ones I have presented as outlined in this paper by Browne et al. and organized in a really nice presentation by Bobak Pezeshki here. These variations involve different ways to explore the tree, expand nodes, simulate the playouts, policies, and backpropation.

And very, very recently Microsoft came out with this exciting paper called "rStar-Math: Small LLMs Can Master Math Reasoningwith Self-Evolved Deep Thinking." The paper shows how a small model could be trained to rival the metrics of the OpenAI o1 model in math. It does this by "exercising deep thinking” through Monte Carlo Tree Search (MCTS)" - once again showing the importance of MCTS in AI.

MCTS with its flexibility, probabilistic nature, and adaptability to domain specific heuristics, excels in both the typical uses of traditional games and atypically in complex real world problems making it an important part of AI innovation and optimizing diverse systems.

Wednesday, January 1, 2025

Machine Leaning Classification: Scikit-Learn, PyTorch, and TensorFlow Examples

Although one can find machine learning examples using scikit-learn, PyTorch, and TensorFlow separately, there aren't really examples where one can see a comparison for these different frameworks on a standard dataset all in one place. But I think it is instructive to see how to use the same variables from the same dataset to accomplish the same prediction task. And that is what we're going to do in this post. We're going to use a very standard diabetes dataset to create basic example classification models based on seven variables to predict if someone is likely to be diabetic. The Github repository is here, the full notebook used to build the models is here, and the deployed Streamlit app can be accessed here.

The seven predictor variables are:

  • Number of pregnancies (the data was trained on all female respondents)
  • Glucose
  • Skin thickness
  • Age
  • BMI
  • Blood Pressure
  • Insulin
The dataset contains an "outcome" column that indicates the presence or absence of diabetes. We'll build four separate models:
  1. scikit-learn (Random Forest)
  2. scikit-learn (Gradient Boost)
  3. PyTorch (Neural Network)
  4. TensorFlow (Neural Network)
The models we will build in this post will focus on basic implementations emphasizing the mechanics and not on other topics like data cleaning, optimization, or fine tuning - although they are also important.

First let's read in the data and create the train/test splits. We'll use the same splits for all four models.

df = pd.read_csv('diabetes.csv')
X = df.drop("Outcome", axis=1)
y = df["Outcome"]
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=78)

Random Forest Classifier:

Create the random forest classifier instance:

rf_model = RandomForestClassifier(n_estimators=100, random_state=78)
Fit the model:

rf_model= rf_model.fit(X_train, y_train)
Make predictions using the testing data:

predictions = rf_model.predict(X_test)

Save the model for the Streamlit app:
filename = 'rf.sav'

pickle.dump(rf_model, open(filename, 'wb'))

Create classification report:
filename = 'rf.sav'
print(classification_report(y_test, predictions))


              precision    recall  f1-score   support

           0       0.80      0.86      0.83       129
           1       0.66      0.56      0.60        63

    accuracy                           0.76       192
   macro avg       0.73      0.71      0.72       192
weighted avg       0.75      0.76      0.75       192

cm = confusion_matrix(y_test, predictions)
disp = ConfusionMatrixDisplay(confusion_matrix=cm)
disp.plot()
As you can see the scikit-learn implementation is pretty straightforward and follows a "model, fit, predict" pattern. Likewise, gradient boost follows the same pattern.

Gradient Boost Classifier:

# Create the gradient boost classifier instance
gb_model = GradientBoostingClassifier(random_state=78)
# Fit the model
gb_model = gb_model.fit(X_train, y_train)
# Make predictions using the testing data
predictions = gb_model.predict(X_test)
# Save the model for the Streamlit app
filename = 'gb.sav'
pickle.dump(gb_model, open(filename, 'wb'))
# Create classification report
print(classification_report(y_test, predictions))

              precision    recall  f1-score   support

           0       0.79      0.87      0.83       129
           1       0.66      0.52      0.58        63

    accuracy                           0.76       192
   macro avg       0.72      0.70      0.71       192
weighted avg       0.75      0.76      0.75       192  

cm = confusion_matrix(y_test, predictions)
disp = ConfusionMatrixDisplay(confusion_matrix=cm)
disp.plot()

Neural Network (PyTorch):

For the neural networks for both PyTorch and TensorFlow, we will scale the data. Scaling the data wasn't really necessary for the tree-based models, but it is for the neural networks. We use the same number of layers and nodes per layer for both the PyTorch and TensorFlow models - 1 feature input layer, 2 hidden layers (16 and 8 nodes with RELU activation functions), and 1 output node that uses a sigmoid activation function. Each of the two networks will run with 100 epochs. And both are using binary cross-entropy for the loss function and Adam for optimization. There is some flexibility of how binary cross-entropy can be set with regards to the sigmoid function between the two frameworks, but how we are doing it here, it is effectively the same.

import torch
import torch.nn as nn
import torch.optim as optim
from sklearn.preprocessing import StandardScaler

Standardize the features:

scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

Now we will convert the data to PyTorch tensors:

X_train_scaled = torch.tensor(X_train_scaled, dtype=torch.float32)
X_test_scaled = torch.tensor(X_test_scaled, dtype=torch.float32)
y_train_tensor = torch.tensor(y_train.values, dtype=torch.float32).unsqueeze(1)
y_test_tensor = torch.tensor(y_test.values, dtype=torch.float32).unsqueeze(1)

Define the neural network model:

class DiabetesPTModel(nn.Module):
    def __init__(self):
        super(DiabetesPTModel, self).__init__()
        self.fc1 = nn.Linear(7, 16)
        self.fc2 = nn.Linear(16, 8)
        self.fc3 = nn.Linear(8, 1)
        self.sigmoid = nn.Sigmoid()

    def forward(self, x):
        x = torch.relu(self.fc1(x))
        x = torch.relu(self.fc2(x))
        x = self.sigmoid(self.fc3(x))
        return x

Initialize the model, loss function, and optimizer:

pt_model = DiabetesPTModel()
criterion = nn.BCELoss()
optimizer = optim.Adam(model.parameters(), lr=0.001)

Train the model:

num_epochs = 100
for epoch in range(num_epochs):
    pt_model.train()
    optimizer.zero_grad()
    outputs = pt_model(X_train_scaled)
    loss = criterion(outputs, y_train_tensor)
    loss.backward()
    optimizer.step()

    if (epoch+1) % 10 == 0:
        print(f'Epoch [{epoch+1}/{num_epochs}], Loss: {loss.item():.4f}')

Evaluate the model:

pt_model.eval()
with torch.no_grad():
    predictions = pt_model(X_test_scaled)
    predictions = predictions.round()
    accuracy = (predictions.eq(y_test_tensor).sum() / float(y_test_tensor.shape[0])).item()
    print(f'Accuracy: {accuracy:.4f}')

Save the model and scaler - we will need them for the Streamlit app:

torch.save(pt_model.state_dict(), 'diabetes_model_pt.pth')
with open('scaler.pkl', 'wb') as f:
    pickle.dump(scaler, f)

Get predictions for test set for classification report:

with torch.no_grad():
    predictions = model(X_test_scaled)
    predictions = predictions.round()

print(classification_report(y_test_predictions, predictions))

              precision    recall  f1-score   support

           0       0.79      0.88      0.83       129
           1       0.67      0.52      0.59        63

accuracy                               0.76       192
macro avg          0.73      0.70      0.71       192
weighted avg       0.75      0.76      0.75       192

cm = confusion_matrix(y_test_tensor, predictions)
disp = ConfusionMatrixDisplay(confusion_matrix=cm)
disp.plot()
Neural Network (TensorFlow):

Define the neural network model:

import tensorflow as tf

tf_model = tf.keras.Sequential([
    tf.keras.layers.Dense(16, activation='relu', input_shape=(7,)),
    tf.keras.layers.Dense(8, activation='relu'),
    tf.keras.layers.Dense(1, activation='sigmoid')
])

Compile the model:

tf_model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])

Train the model. Use the same scaled data used with the PyTorch model:

tf_model.fit(X_train_scaled, y_train, epochs=100, batch_size=32, validation_split=0.2)

Evaluate the model:

loss, accuracy = model.evaluate(X_test_scaled, y_test)
print(f'Accuracy: {accuracy:.4f}')

Save the model for the Streamlit app:

tf_model.save('diabetes_model_tf.h5')

Get predictions for test set for classification report:

predictions = tf_model.predict(X_test_scaled).round()
print(classification_report(y_test, predictions))

              precision    recall  f1-score   support

           0       0.81      0.83      0.82       129
           1       0.63      0.60      0.62        63

    accuracy                           0.76       192
   macro avg       0.72      0.72      0.72       192
weighted avg       0.75      0.76      0.75       192

cm = confusion_matrix(y_test, predictions)
disp = ConfusionMatrixDisplay(confusion_matrix=cm)
disp.plot()
Model Comparison:

These examples show basic implementations of scikit-learn, PyTorch, and TensorFlow models. This is not to demonstrate the huge number of options that are available, such as for fine tuning, data loading in PyTorch, or the properties of tensors in general. We could have also have computed loss curves and done things like adjust for sample imbalance. But from what we did do, we can see the fundamental structure of each of the models.

In comparing the output, we can see confusion matrices that are very similar to each other, which is most likely a function of the data itself or the small size of the data. Normally, these models can vary significantly from each other as far as evaluation of their performance.

Before moving on to deploy these models to a Streamlit app, we could look at one more characteristic of the models, which is to answer the question of which are the most important variables to the model. We'll do that for one of the one of the models - the random forest model.

We can get and display the most important features like this:

importances = rf_model.feature_importances_
importances_sorted = sorted(zip(rf_model.feature_importances_, X.columns), reverse=True)

# Plot the feature importances
features = sorted(zip(X.columns, importances), key = lambda x: x[1])
cols = [f[0] for f in features]
width = [f[1] for f in features]

fig, ax = plt.subplots()

fig.set_size_inches(8,6)
plt.margins(y=0.001)

ax.barh(y=cols, width=width)

plt.show()



As we can see from the chart: glucose, BMI, and age are the most important variables - at least for the random forest model.

Streamlit Application:

For the Streamlit app, we allow the user to enter in values for any of the seven predictor variables and use default values if they don't change them. They can then select from any of the four models we built (and saved) and get a prediction. And very importantly for the two models that we scaled the data, we load the trained scaler for each of those two models and apply it to the user's selections.

And that's it! We have a deployed Streamlit app for the four models.



Sunday, December 22, 2024

Implications of a New Tool for Molecular Dynamics: MDGen

With the end of 2024 and the holidays and extra time off, I've had a bunch of time to catch up on my reading of recent research papers that I've been meaning to get to and one of them that I've been very excited about is a framework called MDGen introduced in a paper called "Generative Modeling of Molecular Dynamics Trajectories" by Jeng et al. that was put out recently. I think this paper needs to get much wider reach, which is one of the reasons I'm writing about it here. But I also want to describe how I believe this system could be extended and how if combined with other systems into a larger pipeline could have a significant impact and really stretch the cutting edge in areas like drug discovery.

Before I get into why I think this paper is important, how MDGen could be used and what its implications are beyond those outlined in the paper, I want to go over a very brief description of molecular dynamics. 

Molecular Dynamics:

Molecular dynamics (MD) plays a pivotal role in drug discovery by providing detailed insights into the movement and interactions of molecules over time. Unlike static structural snapshots, MD simulations capture the dynamic behavior of proteins, ligands, and their complexes, revealing conformational changes, binding pathways, and interaction forces. These insights are critical for understanding how drugs interact with their targets, particularly in complex biological environments where flexibility and motion significantly influence binding affinity and specificity.

In drug discovery, MD is invaluable for identifying binding sites, exploring conformations, and predicting the stability of protein-ligand complexes. By simulating the behavior of molecules at atomic resolution, researchers can assess how candidate drugs bind to their targets, optimize lead compounds, and predict resistance mechanisms. MD also aids in exploring challenging targets like intrinsically disordered proteins, which lack stable structures and require dynamic analysis to uncover potential binding sites. The ability to simulate these dynamic processes accelerates the drug development pipeline, reducing reliance on trial and error methods and enabling more precise, mechanism-driven drug design.

MDGen:

While protein structure predictions like AlphaFold 3, released this past year, has gotten a lot of attention, the problem is that just having that static pose does not get you everything you need, such as building a new enzyme that catalyzes a new reaction. MDGen working with other tools can help with understanding this process.

MDGen is a generative model designed to simulate molecular dynamics (MD) trajectories, offering implications for computational chemistry, biophysics, and AI-driven molecular design. Molecular dynamics simulations, while essential for exploring the behavior of atoms and molecules, are computationally expensive due to the significant disparity between the timescales of integration steps and meaningful molecular phenomena. MDGen addresses this challenge by leveraging deep learning techniques to provide a flexible, multi-task surrogate model capable of tackling diverse downstream tasks.

The generative modeling approach of MDGen diverges from traditional methods, which focus on autoregressive transition density or equilibrium distribution. Instead, MDGen employs end-to-end modeling of complete MD trajectories, enabling applications beyond forward simulation. These include trajectory interpolation (transition path sampling), upsampling molecular dynamics trajectories to capture fast dynamics, and inpainting missing molecular regions for tasks like scaffold design. So this framework expands the scope of MD simulations, making it possible to infer rare molecular events, bridge gaps in trajectory data, and scaffold molecular structures for desired dynamic behaviors.

MDGen is capable of reproducing MD-like outputs for unseen molecules. The model achieves a high degree of accuracy in capturing free energy surfaces, reproducing Markov state fluxes, and predicting torsional relaxation times. Their benchmarks indicate that MDGen can emulate the structural and dynamical content of MD simulations with significant computational efficiency, offering speed-ups of up to 1000x compared to traditional MD methods. Work that was measured in weeks could conceivably be done in hours. This efficiency is particularly advantageous in protein simulation tasks, where MDGen is shown to outperform existing techniques in recovering ensemble statistical properties while being orders of magnitude faster.

With the generative inpainting idea, you can think of this inpainting as like a SORA for molecular dynamics. This inpainting allows for filling in missing molecular regions and generating consistent dynamics for the entire structure. This capability has significant implications for molecular design, particularly in creating new molecules or scaffolding specific dynamics into protein designs. For example, in enzyme engineering, MDGen could generate consistent side-chain configurations and dynamics around a catalytic site, ensuring functional integration into the broader molecular structure.

Not to be too hyperbolic, but I think I'm able to call the implications of this profound, because inpainting inside of trajectories is wild. Because by introducing this generative modeling into MD trajectory data, MDGen enables rapid exploration of molecular dynamics. The framework’s ability to interpolate trajectories suggests a potential for unique hypothesis generation in molecular mechanisms.

Future Possibilites:

Okay, so the potential of the ideas in this paper are incredibly interesting. But I want to outline some of the ideas not specifically mentioned in the paper, how the model could be improved, and potential applications beyond what is readily apparent in the paper.

First, there's the obvious idea of just fine tuning the model to specific use cases. But beyond applying a fine tuned version of the model, it could be worthwhile to look at different tokenization strategies and definitely it would be worthwhile to retrain the model on more and different types of data. For example, it is trained on single proteins. A model could be created using some of the MDGen ideas but trained on protein complexes.

But beyond those types of data and strategies, MDGen could also be trained with multimodal data sources, such as textual or experimental descriptors, which could be very interesting. Furthermore, it could automate experimental design by proposing dynamic behaviors tailored to experimental conditions, guiding laboratory work with predictive insights. Similarly, MDGen could leverage large scale knowledge graphs of molecular interactions and pathways, refining trajectory predictions to include broader biological contexts. These integrations could position MDGen as a versatile tool that bridges the gap between computational predictions and experimental realities.

Furthermore, MDGen's molecular inpainting feature could provide insight into precise design and repair of molecular structures. It could be useful in applications like mutation repair, where it could predict the impact of a mutation and suggest compensatory structural or dynamic changes to restore function. In synthetic biology, MDGen could be used to engineer entirely new molecular pathways with tailored dynamic behaviors, such as light-activated enzymes or thermally sensitive molecular systems. 

The interpolation capabilities could be reimagined as tools for hypothesis generation, allowing the uncovering of unknown intermediate states in biochemical pathways or explore dynamic transitions in materials science. This could significantly aid in understanding complex processes, such as protein-ligand binding or phase transitions at the molecular level.

MDGen could also provide a platform for studying equilibrium and non-equilibrium dynamics, offering insights into phenomena such as protein folding and misfolding in diseases like Alzheimer’s. Its trajectory generation capabilities could be used to explore how time asymmetry manifests at the molecular level, providing theoretical insights into entropy and energy landscapes. As more high-quality MD trajectory data becomes available, MDGen or a model like it that incorporated that data could model increasingly complex systems, offering new ways to study crowded cellular environments or investigate the limitations of Markovian dynamics in highly dynamic systems.

In very practical applications like drug screening and optimization, MDGen could enhance virtual pipelines by predicting dynamic interaction profiles, especially for targets like intrinsically disordered proteins. Its role in multi-scale modeling could bridge atomic-level changes and mesoscopic behaviors, while molecular AI agents built on MDGen could iteratively explore chemical space, design new molecules, and simulate their dynamics for optimized functionality.

Conclusion:

MDGen presents transformative opportunities in molecular science, building on its ability to generate molecular dynamics (MD) trajectories. By framing MD generation as analogous to video modeling, MDGen potentially offers a unified platform for understanding and designing molecular systems. 


Sunday, December 8, 2024

Eigenvalues and Eigenvectors and their Applications

In linear algebra, eigenvalues and eigenvectors are essential concepts in machine learning and artificial intelligence - and are also important in applications in physics, engineering, and much more. They often seem abstract at first, but we can build up an intuition with examples and practical applications.

What are Eigenvalues and Eigenvectors?

 

Suppose we have a square matrix \( A \), which represents some kind of transformation (e.g., rotation, scaling, or shear). When this transformation is applied to a vector \( \mathbf{v} \), sometimes the vector changes direction, and sometimes it doesn’t. When a vector doesn’t change direction (it might still stretch or shrink), that vector is called an eigenvector. The amount by which the vector is stretched or shrunk is called the eigenvalue.
  • Eigenvector \( \mathbf{v} \): A vector that doesn’t change direction when a transformation is applied.
  • Eigenvalue \( \lambda \): A scalar that tells how much the eigenvector is stretched or shrunk.

Mathematically, this is written as:

\[ A \mathbf{v} = \lambda \mathbf{v} \]

Here:
  • A is the transformation matrix.
  • \( \mathbf{v} \) is the eigenvector.
  • \( \lambda \) is the eigenvalue.

A Simple Example: Stretching Along an Axis:


Imagine a transformation matrix: \(A = \begin{bmatrix} 3 & 0 \\ 0 & 2 \end{bmatrix} \) This matrix scales vectors along the x-axis by 3 and along the y-axis by 2. If we apply this transformation to the vector \( \mathbf{v} = \begin{bmatrix} 1 \\ 0 \end{bmatrix} \) , we get:

\( A \mathbf{v} = \begin{bmatrix} 3 & 0 \\ 0 & 2 \end{bmatrix} \begin{bmatrix} 1 \\ 0 \end{bmatrix} = \begin{bmatrix} 3 \\ 0 \end{bmatrix} \)

Here, \( \mathbf{v} \) doesn’t change direction; it’s simply scaled by 3.
Thus:
  • \( \mathbf{v} = \begin{bmatrix} 1 \\ 0 \end{bmatrix} \) is an eigenvector.
  • \( \lambda = 3 \) is the eigenvalue.
Similarly, the vector \( \mathbf{w} = \begin{bmatrix} 0 \\ 1 \end{bmatrix} \) we get:

\( A \mathbf{w} = \begin{bmatrix} 3 & 0 \\ 0 & 2 \end{bmatrix} \begin{bmatrix} 0 \\ 1 \end{bmatrix} = \begin{bmatrix} 0 \\ 2 \end{bmatrix} \)

And here again, \( \mathbf{w} \) doesn’t change direction; it’s simply scaled by 2.
Thus,
  • \( \mathbf{w} = \begin{bmatrix} 0 \\ 1 \end{bmatrix} \) is another eigenvector.
  • \( \lambda = 2 \) is the eigenvalue.

Real World Applications


Before we get more into the math, let's talk about why eigenvalues and eigenvectors are important. Eigenvalues and eigenvectors are more than just mathematical abstractions; they play critical roles in real-world applications. Here are some practical examples:
  1. Machine Learning and Principal Component Analysis (PCA)
    In PCA, we compute the covariance matrix of a dataset and find its eigenvalues and eigenvectors. The eigenvectors represent the principal axes of the data, while the eigenvalues indicate the amount of variance explained along each axis. By selecting the top eigenvectors (those with the largest eigenvalues), we can reduce the dataset’s dimensionality while retaining most of its variance.
  2. Graph Theory
    In network analysis, the eigenvalues and eigenvectors of the adjacency matrix of a graph help us understand its structure. For instance, the largest eigenvalue of the adjacency matrix can indicate the network’s connectivity. Another example is that eigenvectors are used in Google’s PageRank algorithm to rank webpages based on their importance.
  3. Quantum Mechanics
    In quantum mechanics, eigenvalues correspond to measurable quantities like energy levels of a system. The Schrödinger equation involves finding eigenvalues and eigenvectors of the Hamiltonian operator, which gives the possible energy states of a particle.
  4. Image Compression
    In image processing, Singular Value Decomposition (SVD), which is a related concept, relies on eigenvalues and eigenvectors. By keeping the top eigenvalues and their corresponding eigenvectors, we can approximate an image, significantly reducing storage requirements without much loss of quality.

How to Compute Eigenvalues and Eigenvectors:


Alright. Back to the math! How do we actually calculate eigenvalues and eigenvectors?

The determinant plays a critical role in identifying eigenvalues, which leads us to their corresponding eigenvectors.

1. Find Eigenvalues: \( \lambda \):
The eigenvalues of a matrix A are found by solving the characteristic equation utilizing the determinant:

\( \text{det}(A - \lambda I) = 0 \)

And here, \( I \) is the identity matrix \( \begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix} \).

To find \( \lambda \), we solve \( \text{det}(A - \lambda I) = 0 \), which expands into a polynomial equation in \( \lambda \) (called the characteristic polynomial). The roots of this polynomial are the eigenvalues.

Example:

Let:

\( A = \begin{bmatrix} 4 & 2 \\ 1 & 3 \end{bmatrix} \)

The characteristic equation is:

\( \text{det}(A - \lambda I) = \text{det} \begin{bmatrix} 4-\lambda & 2 \\ 1 & 3-\lambda \end{bmatrix} = 0 \)

Expand the determinant:

\( \text{det} = (4-\lambda)(3-\lambda) - (2)(1) \)

\( \text{det} = \lambda^2 - 7\lambda + 10 = 0 \)

Solving this quadratic equation, we get:

\( \lambda = 5, \quad \lambda = 2 \)

Thus, the eigenvalues are \( \lambda = 5 \) and \( \lambda = 2 \).

2. Find Eigenvectors \( \mathbf{v} \):
Once the eigenvalues \( ( \lambda ) \) are identified, we find the corresponding eigenvectors \( ( \mathbf{v} ) \):

For each eigenvalue \( \lambda \), we solve: \( (A - \lambda I) \mathbf{v} = 0 \)

The solution to this equation is a non-zero vector \( \mathbf{v} \).

Example:

For \( A \) and \( \lambda = 5 \) from above:

\( A - 5I = \begin{bmatrix} 4-5 & 2 \\ 1 & 3-5 \end{bmatrix} = \begin{bmatrix} -1 & 2 \\ 1 & -2 \end{bmatrix} \)

Solve:

\( \begin{bmatrix} -1 & 2 \\ 1 & -2 \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} = \begin{bmatrix} 0 \\ 0 \end{bmatrix} \)

This gives:

\( -1x + 2y = 0 \quad \implies \quad y = \frac{1}{2}x \)

Let \( x = 2 \) (arbitrary scaling), then:

\( \mathbf{v_1} = \begin{bmatrix} 2 \\ 1 \end{bmatrix} \)

For \( A \) and \( \lambda = 2 \) from above:

\( A - 2I = \begin{bmatrix} 4-2 & 2 \\ 1 & 3-2 \end{bmatrix} = \begin{bmatrix} 2 & 2 \\ 1 & 1 \end{bmatrix} \)

Solved similarly to find:

\( \mathbf{v_2} = \begin{bmatrix} -1 \\ 1 \end{bmatrix} \)

What does this mean geometrically - and practically?


Given the matrix \( A \) from before:

\( A = \begin{bmatrix} 4 & 2 \\ 1 & 3 \end{bmatrix} \)

This matrix represents a linear transformation in 2D space. When applied to a vector, \( A \) performs a combination of scaling, rotation, and/or shearing.

The eigenvalues and eigenvectors from earlier calculations for matrix \( A \):

  • Eigenvalues: \( \lambda_1 = 5 , \lambda_2 = 2 \)
  • Eigenvectors:
    • Corresponding to \( \lambda_1 = 5 : \mathbf{v}_1 = \begin{bmatrix} 2 \\ 1 \end{bmatrix} \)
    • Corresponding to \( \lambda_2 = 2 : \mathbf{v}_2 = \begin{bmatrix} -1 \\ 1 \end{bmatrix} \)
What These Values Mean
    1. Eigenvalue \( \lambda_1 = 5 \) and Eigenvector \( \mathbf{v}_1 = \begin{bmatrix} 2 \\ 1 \end{bmatrix} \)
    • Geometrically: The eigenvector \( \mathbf{v}_1 = \begin{bmatrix} 2 \\ 1 \end{bmatrix} \) lies along a specific direction in 2D space. When the transformation \( A \) is applied to any vector along this direction, the vector is stretched by a factor of 5, but its direction remains unchanged.
    • Practically: This tells us that there is a “preferred direction” in the transformation where the stretching is maximized by a factor of 5. Vectors aligned with \( \mathbf{v}_1 \) are scaled significantly, making this direction dominant in how \( A \) affects space.
    2. Eigenvalue \( \lambda_2 = 2 \) and Eigenvector \( \mathbf{v}_2 = \begin{bmatrix} -1 \\ 1 \end{bmatrix} \)
    • Geometrically: The eigenvector \( \mathbf{v}_2 = \begin{bmatrix} -1 \\ 1 \end{bmatrix} \) points in another specific direction in 2D space. Vectors along this direction are scaled by a factor of 2 when the transformation \( A \) is applied, with no change in direction.
    • Practically: This direction represents a less dominant feature of the transformation, where stretching occurs to a lesser extent (factor of 2). It shows that the transformation affects different directions differently.
Visually we can imagine a grid of arrows (vectors) in 2D space. Applying the transformation \( A \):
  • Along \( \mathbf{v}_1 \) (eigenvector for \( \lambda_1 = 5 ) \), vectors stretch significantly, growing by a factor of 5.
  • Along \( \mathbf{v}_2 \) (eigenvector for \( \lambda_2 = 2 ) \), vectors stretch less, growing by a factor of 2.
  • For vectors not aligned with \( \mathbf{v}_1 \) or \( \mathbf{v}_2 \), the transformation involves a combination of scaling and changing direction. These vectors can be expressed as combinations of the eigenvectors, showing how any arbitrary vector is affected.

The eigenvalues and eigenvectors of \( A \) provide insights into its transformation behavior. The eigenvectors \( \mathbf{v}_1 \) and \( \mathbf{v}_2 \) represent the preferred directions in 2D space that remain unchanged in orientation under the transformation. The eigenvalues \( \lambda_1 = 5 \) and \( \lambda_2 = 2 \) indicate how much vectors along these directions are stretched or compressed, with \( \mathbf{v}_1 \) experiencing a larger scaling factor.

Furthermore, any vector in space can be expressed as a combination of the eigenvectors, allowing us to decompose the transformation 
\( A \) into its fundamental actions - scaling along these specific directions - providing a complete and intuitive understanding of how \( A \) reshapes space.

Eigenvalues and Eigenvectors in PyTorch


Okay, now that we understand the math, how can we do it programmically in PyTorch?

import torch
# Define the matrix A
A = torch.tensor([[4.0, 2.0], [1.0, 3.0]])

# Compute eigenvalues and eigenvectors
eigenvalues, eigenvectors = torch.linalg.eig(A)

# Display the matrix
print("Matrix A:")
print(A)

# Display eigenvalues
print("\nEigenvalues:")
print(eigenvalues)

# Display eigenvectors
print("\nEigenvectors:")
print(eigenvectors)

# Verify the eigenvalue equation A * v = λ * v for the first eigenvalue and eigenvector
v1 = eigenvectors[:, 0] # First eigenvector
lambda1 = eigenvalues[0] # First eigenvalue

verification = torch.matmul(A, v1) # A * v
expected = lambda1 * v1 # λ * v

print("\nVerification for the first eigenvalue and eigenvector:")
print("A * v1 =", verification)
print("λ1 * v1 =", expected)


The output will look like this:

Matrix A:
tensor([[4., 2.], [1., 3.]])

Eigenvalues:
tensor([5.+0.j, 2.+0.j])

Eigenvectors:
tensor([[ 0.8944, -0.7071], [ 0.4472, 0.7071]])

Verification for the first eigenvalue and eigenvector:
A * v1 = tensor([4.4721+0.j, 2.2361+0.j])
λ1 * v1 = tensor([4.4721+0.j, 2.2361+0.j])


Now if you have been following along closely, you may be wondering why the eigenvalues from PyTorch of 5 and 2 match what we calcuated manually, but the eigenvectors \( \mathbf{v}_1 = \begin{bmatrix} 2 \\ 1 \end{bmatrix} \) and \( \mathbf{v}_2 = \begin{bmatrix} -1 \\ 1 \end{bmatrix} \) do not match \( \mathbf{v}_1 = \begin{bmatrix} 0.8944 \\ 0.4472 \end{bmatrix} \) and \( \mathbf{v}_2 = \begin{bmatrix} -0.7071 \\ 0.7071 \end{bmatrix} \).

This is because when solving the eigenvector equation \( (A - \lambda I)\mathbf{v} = 0 \), we calculate the eigenvectors up to any scalar multiple. For example, if [2, 1] is a solution, then [4, 2] or even [0.2, 0.1] are also valid eigenvectors because eigenvectors are defined up to scaling. And if you look above when we were calculating eigenvectors, we let \( x = 2 \) (arbitrary scaling).

PyTorch automatically normalizes the eigenvectors so that their length (or magnitude) is 1. This is achieved by dividing each eigenvector by its norm:

\( \|\mathbf{v}\| = \sqrt{x_1^2 + x_2^2} \)

For example:

• For the manually calculated eigenvector [2, 1]:

\( \|\mathbf{v}_1\| = \sqrt{2^2 + 1^2} = \sqrt{5} \)

Normalizing it:

\( \mathbf{v}_1 = \left[\frac{2}{\sqrt{5}}, \frac{1}{\sqrt{5}}\right] = [0.8944, 0.4472] \)

• Similarly, for [-1, 1]:

\( \|\mathbf{v}_2\| = \sqrt{(-1)^2 + 1^2} = \sqrt{2} \)

Normalizing it:

\( \mathbf{v}_2 = \left[\frac{-1}{\sqrt{2}}, \frac{1}{\sqrt{2}}\right] = [-0.7071, 0.7071] \)

Both the manually calculated eigenvectors ([2, 1] and [-1, 1]) and PyTorch’s normalized eigenvectors ([0.8944, 0.4472] and [-0.7071, 0.7071]) are equally valid. They represent the same direction in space, and the eigenvalues (5 and 2) remain unchanged. The difference is only due to normalization. This ensures a standardized representation of eigenvectors, which is particularly useful in programming and numerical computations. Many applications, such as Principal Component Analysis (PCA), assume normalized eigenvectors for interpretation or computation.

Summary

Eigenvalues and eigenvectors provide a great way to understand transformations in linear algebra. They appear in diverse fields, from machine learning and physics to structural engineering and image processing. We can use these tools to solve complex problems and gain insights into the behavior of systems.



Tuesday, November 19, 2024

The Evolution of Market and Political Research: AI Agents

The landscape of market and political research is about to undergo a significant transformation. Traditional methods like surveys, panels, and in-person focus groups, while long-standing, will increasingly be replaced by AI-driven alternatives. These methods, leveraging autonomous agents to simulate human behavior and attitudes, are proving to be faster, more cost-effective, and potentially more accurate. This shift represents a turning point in how we gather insights, and it is happening much sooner than many anticipated.

There are two papers that have recently come out that I want to use to illustrate what I believe is going to be possible:

  •  "Generative Agent Simulations of 1,000 People" is a paper that just came out of Stanford. The paper presents an architecture for simulating human behavior using generative agents informed by qualitative interviews. What's really interesting is that these agents are modeled after 1,052 real individuals. They replicate attitudes and behaviors across various social science tasks with high accuracy, performing comparably to human self-replication over time. The research demonstrates the agents’ utility in predicting responses to surveys, personality assessments, economic games, and experimental settings. By reducing demographic biases and enabling scalable simulations, the approach offers a powerful tool for understanding individual and collective behaviors in diverse contexts.
  • "Scaling Synthetic Data Creation with 1,000,000,000 Personas" is a paper that introduces Persona Hub, a collection of one billion synthetic personas designed to enhance data diversity and scalability in large language models (LLMs). By associating each persona with unique perspectives and knowledge, the framework enables the generation of highly diverse synthetic data across multiple applications, including math problems, instructions, and knowledge-rich texts. The approach overcomes limitations of previous data synthesis methods by leveraging personas to guide LLMs, demonstrating significant potential for advancing AI research, development, and practical applications

These are just two examples of the recent creation of agents, but there are many others and I have also been creating my own autonomous agents that can do focus groups and questionnaire research. So I strongly believe that AI and agents will have a large role in the future. But before we can talk about using AI agents, we need to examine the limitations of traditional research.

Limitations of Traditional Methods

I started out in market research many years ago in a part-time job during college checking data quality in survey questionnaires. When I graduated, I worked as an analyst for a company called Sophisticated Data Research (SDR). I left SDR after a few years, dissatisfied with the current state of software at the time for research and went on to join another company to write some of the first statistical software for Windows and for the internet for market research. In the early 2000s, I left to join a start up to build agents to model marketing effectiveness, so I've been around agents for over 20 years. So I'm very aware of agents, but also of the traditionalism in the industry and its limitations. 

Traditional approaches to market and political research have long faced challenges, but these issues have become more pronounced in recent years. Phone-based surveys, once a cornerstone of consumer and political research, have seen their accuracy steadily decline. The widespread use of mobile devices has fundamentally changed how people interact with calls - screening is common, response rates have plummeted, and the pool of reachable participants is increasingly skewed. This has led researchers to rely on heavy weighting of subpopulations to align with presumed demographic truths. However, this practice has become increasingly tenuous, bordering on speculative guesswork, as the assumptions underlying these adjustments often lack a solid foundation. As a result, the reliability of phone surveys is now widely questioned, making them an increasingly impractical method for gathering actionable data.

These challenges extend beyond phone surveys. In-person focus groups and panels also face issues with scalability, cost, and bias. Facilitating these sessions requires significant resources, and their relatively small sample sizes make it difficult to generalize findings. Biases - both from facilitators and participants - can further distort results. Focus groups are increasingly having a difficult time recruiting in some segments - doctors, researchers, engineers, etc. Together, these factors have created a pressing need for new methodologies that are more efficient and reliable.

The Role of AI and Agents in Simulated Research

Recent advancements in artificial intelligence, particularly in the creation and deployment of generative agents, are addressing many of these challenges. By using large libraries of personas, AI systems can simulate the attitudes, preferences, and behaviors of diverse populations - and difficult to reach populations like doctors. Studies have demonstrated that these agents can replicate human responses with a high degree of accuracy. For example, as mentioned earlier, research from Tencent’s Persona Hub highlights the ability to synthesize billions of personas, enabling nuanced and scalable simulations, while Stanford’s work on generative agents shows their effectiveness in predicting individual attitudes and behaviors.

These systems allow for the creation of virtual focus groups and the simulation of surveys in which each participant is an AI-driven persona. In virtual focus groups, the personas can interact dynamically, mimicking the complex interpersonal dynamics found in real-life settings. These approaches don't have to wait to recruit participants or field a study. They can be done immediately and not just done once but repeated hundreds or thousands of times. This approach enables the collection of insights that are not only faster to obtain but also potentially more comprehensive.

Benefits and Broader Implications

Simulated research offers several advantages that are increasingly difficult to ignore. First, it reduces the time and costs associated with traditional methods. Virtual focus groups can be conducted instantaneously and at a fraction of the cost, making it possible to run studies that were previously too expensive or logistically complex.

Second, the accuracy of these methods is rapidly improving. Generative agents have demonstrated their ability to align closely with human responses in studies, offering reliable insights that rival or exceed those obtained through traditional research. This capability challenges the reliance on demographic sampling by using more detailed persona-based approaches, which can reduce biases.

Third, the technology is evolving to overcome limitations such as knowledge cutoffs in large language models. Although, not too many people are talking about this, but in a paper titled "Mixture of a Million Experts" the authors talk about this idea of continuous learning. And this is really exciting - AI models are going to be able to continuously be updated. Continuous learning capabilities will enable AI agents to stay updated with real-time information without relying on web searches. This development will further enhance the utility of simulated research by providing more contextually relevant and up-to-date responses. This will enable the type of research based on recent current events - opening up a potentially effective tool to measure practically real-time economic and political attitudes.

Preparing for the Future

The rise of AI-driven research methods signals a need for companies in the market and political research sectors to rethink their approaches. Adapting to this new reality will require investing in AI capabilities and integrating them into existing workflows. Organizations will also need to reconsider their business models, as the cost structures of traditional methods are unlikely to remain competitive against the efficiency of synthetic research.

Some of the larger organizations will be unable or unwilling to adapt as they try to protect the ways they have been doing things going back decades. Most will try and add AI to their current offerings as a sincere but ultimately half-hearted attempt to remain relevant. They will talk about things like their agent based profiles in their new AI based consumer segments. First, if you are talking about your new "AI-based insights" as part of your new marketing, well everyone is saying that now and how is that any different than saying in the early 2000's that your new solution is using the World Wide Web. How does that excite a customer - when you are stating something obvious and what everyone else is saying? Second, don't just tack on AI onto your existing offerings. That's not  going to fly in a time of exponential change. In order to adapt to exponential change, you need to think radically. 

Because with the coming of agents, there will be some use cases where the cost of doing research will be driven to zero and the barrier to entry will be minimal.

While this transition may not render traditional methods entirely obsolete overnight, it is clear that the trajectory of research is changing. The industry must embrace these advancements to stay relevant in a world where insights will become increasingly instantaneous and accessible. If old value propositions are driven to zero, new value propositions and strategic advantages will need to be identified.

An Inevitable Shift

The adoption of AI-driven research is not a distant prospect; it is already happening. As the tools and techniques improve, they will become integral to understanding consumer and voter behavior. The question for organizations is not whether to adopt these methods but how quickly they can do so and how effectively they can integrate them into their operations.

The AI transformation of market and political research signals that innovation doesn't merely enhance - it redefines and disrupts industries entirely. AI agents are not just an alternative to traditional methods - they are a glimpse into the future of how we understand and engage with the world.



The Cache Is the Thought — What KV Caching Reveals About How AI Actually Works Machine Intelligence · Technical Essays Apri...