Showing posts with label Assembly. Show all posts
Showing posts with label Assembly. Show all posts

Sunday, September 6, 2026

SUDO INIT 6: Rebooting Your C Journey

 

                                                    Yeah, let's do just that. So grab some popcorn and lean back.

 

I've said it again and again: To teach is to lie.

 

I'll simplify things. A lot.

Not that the details aren't interesting, but if I start throwing around PTEs, PFNs, page faults, kernel internals and CPU privileges from the beginning, nobody is going to read this.

Worry slightly, though. I'll throw some stuff your way. End of warning.

 

I Wanted a C project


I've been learning C for a while and I genuinely like it. I like it a lot. But I don't like the "just build a project!" approach to learning.

Build an app.

A website.

A to-do list.

A game or something that simplifies your life.

 

Just do a project = me dead inside

 

Yeah, that's not me. I'm not a programmer. I don't want to build an app and spend months adding features, fixing bugs and maintaining it. That's not me. To me, these kinds of projects are mood-killers.
 

What I do enjoy is understanding what the machine is actually doing. I want to use C to ask questions about the stuff I really want to understand.

And I've been having some success learning C and Assembly while also playing around with microcontrollers and lower-level concepts: memory, pointers, addresses, system calls, how OSes work, and so on. 

 

A Memorable Idea?

 

 I knew that a process doesn't deal directly with physical RAM addresses. It works with virtual addresses.

 

The CPU and OS cooperate to translate those virtual addresses into physical addresses. The OS sets up the mappings; the CPU's memory-management hardware uses them.

The details are considerably more interesting, but let's keep this mental picture for now:

Process --> Virtual Address --> Translation --> Physical Address 

 

I also knew that, under certain conditions, the relationship between virtual and physical addresses could change.

The same virtual address could end up referring to a different physical page.

So, could I force those conditions? Can I make one of those addresses change?

Hey, isn't this a project of sorts?

 

I was hyped. And that's a good sign. If you are ecstatic about doing something that is useful for what you're studying, something you actually want to become good at, that's a good sign. 

 

The Plan

 

 I started by asking myself some questions and starting small:

- Can I get the PID of a process? Of course, easy.

- Can I allocate a page? Probably.

- Can I print the virtual address? Sure. 

- Can I figure out which page that address belongs to?

- Can I find the corresponding physical page?

- Do I have enough privileges for anything related to the physical address?

 

And finally, when I know I can access all of this and print it to stdout:

- Can I make the relationship change while the process is still alive? 

 

Oh, but for context mmap() in ASM was the touchstone that made me trip into this rabbit hole. I even considered building my whole project in ASM, but then briefly considered how much pain I was willing to accept. So C it was!

mmap()


If you haven't used it before, mmap() is a Linux/POSIX interface that lets a program create mappings in its virtual address space.

 

For my first experiment, I used it to create a mapping for a single 4096-byte page of memory. Remember that 4096 bytes in decimal is 0x1000 in hexadecimal. 

 

This was important because 4096 bytes is also the page size I'm working with here.

A 4096-byte page needs 12 bits to identify any byte within it, because:

4096 = 2¹²

So, the lowest 12 bits of a virtual address tell us the offset within the page. 

Everything above those 12 bits gives us the virtual page number.


Once we have the virtual page number, we can find the corresponding entry in pagemap

There is one 64-bit entry in pagemap for each virtual page. So, to find the entry for our virtual page, we multiply the virtual page number by 8. 

 

A pagemap entry consists of several pieces of information. I care here about the Page Frame Number (PFN), which tells me which physical page is backing the virtual page. 

Then I wrote a value into that memory:

42 

 

So far, nothing particularly exciting. Just moving along and gathering our LEGO bricks.

 

But now we had our program, a process, a virtual address, a page, and a value stored at that address.

I made the program request the user to press the Enter key. This way, it was kept alive.

 

But where is that page actually sitting in physical memory? 

Linux provides an interesting interface for this:

/proc/<PID>/pagemap

 

This is a kernel-provided interface that lets a userspace program inspect information about a process's page mappings.

In particular, a pagemap entry can contain the PFN, corresponding to a virtual page.

And, from here, the physical address is fairly straightforward:

Physical Address = PFN x page size + page offset

  

So, now I had all the bricks that I needed.

I also learned that you needed the appropriate privileges to actually access the pagemap

Root isn't necessarily equivalent to having every Linux capability. On my system, I needed the privileges associated with CAP_SYS_ADMIN to obtain the PFN information from pagemap.

And that gave me another little rabbit hole: being root and having a particular Linux capability aren't quite the same thing.


 

 

 

 

 

 

 

 

  

 

 

 

 

 


Experiment A (first steps) or, as  Ilike to call it: To Sudo or Not to Sudo 

 

The Experiments

 

I called these the experiments. Gathering all the tools I needed was Experiment A. I now wanted to jump into Experiment B, where I'd try to change the physical address.

I knew one cheaty way to do it.

I knew about Copy-On-Write (COW). I'd recently been reading about COW in another context, including some security research involving COW-related behaviour.

So this seemed like a natural place to start.

 

The basic idea behind COW is pretty simple to understand:

Suppose a process has a page containing some data. Now we create a child process with fork() - see here.
The child gets its own process and its own virtual address space. 

But, for the moment, the kernel has no need to immediately make a complete physical copy of every page.

Instead, parent and child can initially refer to the same physical page.

The mappings are made effectively read-only (this is important) and the system keeps track of the fact that the page can be copied if someone tries to modify it.

  

Experiment B

 

 

 Why 666? Because Gilfoyle is my spirit animal

 

Now, my child process did the following:

*(int *)addr = 666;


The CPU detects that the attempted write isn't permitted by the current page-table entry and raises a page fault.

Now, before you start thinking segfault, nope. 

A page fault does not automatically mean something has gone wrong.

A page fault is a mechanism, and the kernel can handle it.

In this case, it recognizes that this is a legitimate COW situation, creates a new physical page for the child, copies the contents, changes the child's mapping, and allows the write to continue.

 

And, voilà! Same virtual address in two processes, different physical addresses.

So, slight cheating, but not entirely.

 

Also, see the weird thing up there? Maybe you don't. Compare that image of experiment B against my running that exact same program a second time:

 

 See it now? Both processes are writing to stdout and we cannot predict which one will print which line first. 

 

For a moment, I was baffled by this, but then I remembered what fork() actually does. It's not instructing the child to run after the parent runs. It means:

There are now two concurrent processes capable of running. Run them.

The scheduler decides when each gets CPU time.

So, voilà! An accidental demonstration of concurrency. You're welcome. It wasn't planned.

As for concurrency, that's another rabbit hole. Rabbit holes everywhere you look. Reminds me of math classes and the Real line. My professor used to draw rabbit holes there all the time. Fun, fun, fun.


 Experiment C

 

 Single process, same physical page, two different virtual addresses.

I now wanted to remove fork() from the equation. 

 This required a different approach, and I used a POSIX shared-memory object and mapped it twice into the same process.


Think of the shared-memory object as the thing being mapped. The mapping itself isn't the memory.

I asked Linux to map the same object twice.

So, now Linux could give me two different virtual addresses while both mappings referred to the same physical page. 

 

  

 

 

 

 

 

 

 



Note: it is not a given that the two addresses will be adjacent. 

 

I could inspect both virtual pages through /proc/<PID>/pagemap, once more and check that the PFN was the same and, therefore, the same physical page.

 

The Issue

 

I had technically answered part of my original question, but I had used fork(). What if I didn't?

Could I get a single process, keep the same virtual address, and somehow make Linux replace the physical page underneath it?

That became the next question. 

And this is where the project is currently sitting.

I already received a very interesting suggestion for how to approach it, involving madvise() and memory pressure.

 

I haven't properly investigated it yet, so I'm not going to pretend I understand it well enough to explain that here.

That's the next rabbit hole, and one you might want to trip into yourself. All quite accidentally, of course.

 

 

 

 

 

 

 

 

  Yup, I'm a cute little sunrise.  

 

 

 

 

 

 

 

I started this little field trip because I wanted to learn C and understand virtual memory better. And, along the way, I got to touch on:

- mmap()

- pointers and pointer casts

- virtual pages and page offsets

- PFNs

- /proc/<PID>/pagemap

- Linux capabilities

- COW

- page faults

- shared memory

- multiple virtual mappings

- process IDs

- scheduling and concurrent processes

- ... 

 

If anyone reads this and is interested in its innards, here's my Github. Mind the step, not the mess.

 

The code is deliberately small, and certainly not a production-quality memory inspection tool.

It's an experiment. Or a series of them, actually.

Something you can compile yourself, run, modify, break, and have fun with. 

 

One last thing:

The little snippets of these experiments that I posted on LinkedIn were well received. People have suggested alternative experiments, pointed out things I wasn't considering, asked interesting questions which themselves gave me ideas of stuff to try.

That's been rather nice.

I don't pretend to be an expert in any of these things. Quite the contrary. So I'm humbled and grateful when those that know more than I do chime in and give their input.

 

I was going to say keep digging, but that's not necessary. Just accidentally trip into any of the rabbit holes that are all around us. And be perpetually amazed as you fall deeper and deeper. 

 

 

Saturday, October 11, 2025

How Computers Think: A Liars Walk Through x86 Assembly

 

                                                                                    So, yeah. The artist found out about digital art.
 
 

"To teach is to lie."

 

A few days ago I promised to talk a bit about a simple Assembly program.

And talk I shall!

First off, let's just get out of the way what that code is doing:

It takes any number of integer positive values as parameters and adds them.

So, if you run the program with ./<program_name> 40 2 (which we shall in a bit) it will print out the answer:



If this were a Python program, it would be simple as can be. But Python is withholding from you a lot of what is happening behind the scenes. And we want none of that. We want to see stuff as it is happening.

 

First of all, a couple of disclaimers:

- This code isn't as simple as it could be. It's in fact part of a series of lessons I'm going through in order to review and solidify my ASM and debugging knowledge.  They are part of a NASM tutorial (ASMtutor), which I highly advise — it's quite carefully made, without major errors (very common in basic ASM stuff) and which guides you in a smooth manner towards greater knowledge. You can find it here.
- The code I'm using can be found here.

- To create a compiled binary of this code, in particular of the files functions_v1.asm and 13_atoi.asm, you need to download those two files and run: 

nasm -f elf 13_atoi.asm

ld -m elf_i386 13_atoi.o -o atoi

 This will create an executable named atoi, which you can then run as explained before.

- You should know a bit about binary and hexadecimal. Please take a quick gander at this. Knowing how to quickly translate between them will really help to inspect the debugger.
- And speaking of which, I'll be using GDB with GEF and I advise you to do so as well. If you already rock on GDB you probably don't need it. But then again, I highly doubt you'll be reading this if you already have mastered GDB. 

As for GEF, it will make the debugger much more pleasing to the eyes, and your code easier to follow.
- To finally start the debugger (GEF) with our program, we'll run:



- My focus isn't a 'deep dive' into Assembly. I want to give you the very basic and essential tools to be able to go through this code, instruction by instruction and be able to understand how the computer is 'thinking'. In order to do that, I'll give you some basic tools.

- Speaking of which, here is a pdf with the few commands and concepts you'll refer to in order to be able to follow through the code.
- Functions in Assembly are interesting. You can call a function or jump into a function (check the PDF), but also can just enter a function because you're going through the code flow and there is no call, no jump (or there isn't a reason to jump). So, for example, in this case:

You have here a comparison and you are checking if eax is equal to 0, when you do cmp eax, 0. Then you have a jump not zero. If eax is NOT zero, then you jump elsewhere to the function divideLoop. But if eax IS in fact zero, then you just move along that grocery shopping list to the next item on the list. Therefore, you enter the function printLoop and decrease ecx by 1. 

 

On the other hand if printLoop actually had a dot behind it (.printLoop) it would be a function inside of whatever function we are in, and couldn't be called from outside of that function. If in doubt, dig a little deeper into Assembly. 

- You might also need a quick contextual/theoretical rundown, so here it goes:

Assembly is the lowest level code you can have that is still human-readable. Below it is the world of ones and zeroes, and above it the world of different programming languages and their way to read and write programs.

In assembly, we're working in the so-called User Space, but sometimes we need to give control to the Operating System through system calls. To do that, we use interrupts (Here is a link to a nice Linux system call table. If you use Windows we can't be friends, sorry). You won't really need that table in order to follow along and read the code, but it sure helps. We only use two syscalls in this code — one for writing to the terminal, and another to exit cleanly. (look up: do the registers change after the OS is done with them?).

When the syscall happens, the OS takes control of the program, and uses the stack and your registers. After that, you cannot be 100% sure of how all registers will be, so if it matters to you, like what happens in other situations, make sure you save those registers into the stack and get them back after the OS lets go of control and returns it to User space.

 

There are a couple of bitwise operations - basically XOR ('exclusive or'). To learn more about bitwise operators, check this page. And if you like computer games and are even slightly interested in electronics or logic, then take a gander at this game. It's great fun.

But, again, you don't have to jump into bitwise operations to read this code. Just do a mental substitution. Whenever you see, for example xor eax, eax it's basically the same as taking zero and placing it in the register eax or, in other words, doing mov eax, 0

Learn boolean logic, bitwise operations, and everything else because you want to and because its fun. :) 

 

What is a register? What is the Stack? You probably heard about those two terms before. Here's a very quick and dirty analogy:

Imagine you're sitting at a desk and I'm telling you to remember things, like a name, then a number, then an address, then another name. Sure, you can do it for a while. Your short term memory will be able to memorize some stuff before it's too much stuff to handle. Those places where you're storing this easy-to-remember, short-term memory stuff are the registers. Registers are temporary containers for fast memory access. Incredibly fast memory access. A bit like short-term variables on steroids.

But what if I start giving you too much stuff to remember? Then you are allowed to write whatever information I tell you on small pieces of paper, as long as you store them properly on a receipt-spike (also known as an order-spike). You've seen them before in restaurants. Something like this:
 

                          Analogies R Us: each strip of paper a memory location, an integer, etc
 
This stack, or this order-spike, uses a LIFO system, meaning Last In First Out. Whatever last piece of paper with information you have pushed into that pile, is the last one you can pop out and examine.

Neat, no? So, you can store x amount of things for very quick access inside those registers, and you can push whatever you can't hold in quick access memory into that pile, and then you can take it our and again place it in a register if you so wish (actually, you push copies of those values — the originals remain in registers unless you clear them).

All that I've said so far is to be taken with a pinch of salt. Not only because these last bits are an analogy, but because, like I said in the beginning, to teach is to lie. I have to omit stuff in order to let you do some progress.

Imagine us teaching Math in class and instead of telling kids all over the world that you can't divide by zero, we instead said, there are several instances in mathematics where dividing by zero is perfectly sensible, and then described said instances. Can you imagine the faces of those 7th graders? Nothing would be retained.

By lying to them, they learn a useful rule and important aspect of mathematics, and when they advance in their mathematical careers, they get to learn at least of one way to actually divide by zero.

It's the same here. I'm skimming a lot of fat in order to draw an understandable picture, and to let you (if you so wish) take your first steps in Assembly, which I truly love.

Got any questions? Stuff you don't understand? Stuff you'd like to see? Shoot them in my general direction. Search on Google, use an LLM ("gulp, LLMs?! But aren't they the devil?". Look, I was here when Wikipedia first came up and academia was having a fit over it. I also survived using AltaVista back in the 90s. You'll do fine with LLMs and live to tell the tale).

As I've shown in one of the pictures above, we'll be using our program to add 40 to 2 and to find out the Answer to the Ultimate Question of Life, the Universe, and Everything.

I won't go through the whole code. It would serve little purpose and, besides, you need to walk the walk.


                                                     Ah, Jayne... you look young.
 

 So, back to business!

 

You need to:

- download the two .asm files;

- compile them into a binary file;

- install GEF (GDB probably already installed);

- download the PDF with some information;

- Run GDB and check what happens to the Registers and to the Stack as you step into the program.

 

And that's it! We got a party going.

 

Ok, let's pretend anyone is actually reading this and actually is inspecting the program with the debugger. You run the code in terminal to open the debugger, then when it opens, you write start and press enter.
What might you be seeing at this point in time? Something like this:


                                        Already regretting your very recent life decisions? Don't! This stuff is fascinating.


What a bloody mess! I know. Don't worry. The first time I started looking at this I was pretty confused as well. But let's first try to get some 'unhelpful' information out of the way (ahem, teaching is lying) and focus on the meat of this and at what we'll be focusing on exclusively. So, clean-up time!

 

                                       Some breathing room, at last!
 

Aha, so what do we have here?

We have 3 different screens:

- Registers

- Stack

- Code

 

Lucky us! Because that's exactly what we want to follow as we're moving along our code. Let's start with the registers. If you've looked at your PDF, you'll see the stars of our show there. 

Register Screen:

esp is the stack pointer and it is always pointing to the top of the stack. It's not holding the value at the top of the stack, which is the value 3, but it's holding the value 0xffffd0f0.

Huh.. interesting. If you look at the menu below, you'll see that that's the exact stack address of its topmost value. Memory position 0xffffd0f0 is the topmost piece of paper (like its identifier) and in it we have written the value 3 (more on why 3 later). Everything else we care about is (apparently) 0 for now, so let's move on.

Stack Screen:

Remember that pile of papers stuck at the order-spike? This is it.
 Here's an example of one of these lines:

0xffffd0f8│+0x0008: 0xffffd2e0  →  0x32003034 ("40"?)

Let me explain it to you. 

0xffffd0f8  memory location in the Stack. This is represented in 4 byte intervals, since we're in a 32 bit program

+0x0008  Memory offset from the top of the stack. This means that we're in a position 8 bytes above the position of the top of the stack (0xffffd0f0). As we add stuff to the stack the memory values go down. This is an important quirk that you would do good in remembering. Also note that while this has the stack facing upwards, sometimes you'll see it upside down (which makes the memory values 'growing' down appear to be a bit less strange).

0xffffd2e0  What is stored in that stack position (our piece of paper). In this case it is a pointer to a string. That string is '40'. So, in other words, there is a location in memory that is storing the string '40' and in this stack there is a pointer pointing to the first character of that string, or '4'. Read on, it will be made somewhat clearer. You don't need to know in depth what are pointers, etc to be able to read this program, but knowing these things de-mystifies them plenty and makes your life easier and more enjoyable. 

(bonus cookie points for you if you can explain why, as you go up in the stack you go from  0xffffd2e3 to 0xffffd2e0, a small 3 byte jump, and then follow with a big big jump from 0xffffd2e0 to 0xffffd2a5 a 59 byte-sized jump)

0x32003034  Well, I see a 40 and the beginning of a 2. :) Let me explain:

As 40 and 2 were used as arguments in our program, they were added to the stack. But they weren't added as integers. In fact, they were added as ASCII characters. 


 We're working with hexadecimal characters and we're looking for the value 40, right?

So, 0x32003034. You'd possibly expect to read from left to right, but we're working in little-endian here, so in fact our number will be presented right to left. And if we look at that table, we would be reading (now reversed) 4, 0, NUL, 2. That NUL is the null byte which marks the end of that string. And I'd bet you my breeches that after that 32 there will be immediately to its left another 00, since these strings are being stored next to each other. If you're paying attention, you'll know I was cheating and you won't want to bet against me as you read the next value in the stack 0x48530032 and see, in black and white that '00 32', which translates to the number 2 (in decimal). You might be tempted to continue translating that line and read 'SH..'. Interesting, no? But I'll leave it to you to find out more.

But you might be asking, why is the Stack in this initial state? And that is a perfectly valid question. As we load up our program and enter its two parameters 40 and 2, we have automatically added 4 things to the stack, in this order (from bottom to top):

- a pointer to the last parameter (our third argument, or argv[2]);

- a pointer to the first parameter (our second argument, or argv[1]);

- a pointer to the program itself, its full path, in fact (our first argument, or argv[0]);

- the number of arguments  (argc).

As you start reading the program and debugging it, you'll see the stack add more items and remove them. 

Code Screen:

This is where you follow your code. the green arrow and green letters shows you where you're at in the code, and the red dot tells you that there's a breakpoint here. If you're into learning more about GDB (please do) then you can try this link as well. This line of code is the next to be executed. And as you press si ('step into') you'll move along the code, Assembly line by Assembly line.

 

So, let's do just that and run si some 4 times, until we get to that compare (cmp). How will that code look like? That Stack, will it change? And the Registers?  Let's see:


 So, what happened? We started by popping the value on the top of the stack and placing it in ecx. So, in that next step you'll see that ecx = 3. Then we do the same and place that value, a pointer to the function name, in edx, removing it from the top of the stack. Finally, we decrease ecx, basically subtracting 1 to the value within and we xor edx, edx which is basically the same as doing mov edx, 0. ecx is holding the number of variables used in our Math and edx is emptied for future use. See how that works? If we push, we take the value in the register and copy it into the stack, and if we pop we remove the value from the stack and place it in the designated variable.

I'm stopping here with the guided code lesson. I'm leaving it up to you. You do have all the tools that you need to do so.

Remember tidbits like: if you call a function, you jump to that function, instead of continuing down the lines of code, but you place the line of code right after the call on top of the stack (the return address is pushed onto the stack). When you hit the instruction ret, on the other hand, you pop whatever value is on top of the stack and go directly to the line that is referenced in that value (a pointer to a memory value). 

GEF will be kind enough to tell you when a jump instruction isn't fulfilled (which means that instead of jumping into a function you continue reading the code below that jump).

Pay also attention to the name of the function and on how far you are away from its start. For example, in this case, I'm inside the inner function .multiplyLoop, inside the function atoi:

 


 

Here's some more stuff you might want to get into:
 

- what is the register eip doing?

-  What's with bl, BYTE PTR [esi+ecx*1]? Can you guess what's happening there?

- Pay attention to where you're going and what's happening when you run a ret instruction.

- Perhaps write down all the commands in a piece of paper and try to 'guess' how the flow should work until the very last operation. Then compare to what you see on GDB. Does it match?

And that's it! You're now ready to start discovering more and more about this wonderful world, or you're thinking that anyone that likes this must be a bit crazy.



Either way, you cannot unsee it now. 
Have fun and hack away!

Saturday, July 26, 2025

"INTs Aren't Integers and FLOATs aren't Real"

 

                            I was told this is a cat-submarine. Tail = Periscope. I believe it.
 


Over the past few weeks, I’ve been juggling two realities. On one side work: networks, and the daily exploration of security tasks in a banking environment. On the other, low-level code: free time, NASM, one instruction at a time.

 

I’ve been reviewing the basics—x86 syntax, memory layout, data definitions, etc etc etc. I'm following a series of videos, from a YouTuber that I appreciate, and although I found some information lacking or imprecise—particularly in the episode devoted to divisionit's still a good set of videos and I'd advise anyone to watch them here. Lately I've taken a gander at integers. How they’re stored, manipulated, compared, and what all those flags mean when you're moving bits around and trying to make sense of a program in GDB.

If you’ve spent more than a few hours in GDB (as I have, unreasonably so at times), you’ve probably done a CMP eax, ebx and then wondered what exactly happened to the flags. What’s the deal with CF, ZF, SF, and the rest of the alphabet soup? Why do certain jump instructions follow CMP, and not others?

So, quick refresher for the two family members that like me and still read this blog:

CMP just subtracts the second operand from the first, sets flags, and discards the result. The flags tell you the outcome, and then you pick the jump based on what you want to test.

InstructionMeaningFlags checked
JE / JZEqual (zero result)ZF = 1
JNE / JNZNot equalZF = 0
JL / JNGELess than (signed)SF != OF
JLE / JNGLess or equal (signed)ZF = 1 or SF != OF
JB / JCBelow (unsigned)CF = 1
JAAbove (unsigned)CF = 0 && ZF = 0

Notice those signed vs unsigned differences? JB isn’t "jump if smaller", it’s "jump if below"—unsigned. If you’re comparing signed ints, you should be using JL, JG, and so on.

Floats make it even more nuanced. UCOMISS xmm0, xmm1 for example, is how you compare scalar floats. That instruction sets flags similar to CMP, but works with IEEE 754 single-precision values, not integers. And yes, it’s aware of signs, NaNs, Infs, and the rest (of floating point hell).

Anyway, all that to say: this is kinda subtle, or at least it requires some study and care. IMO, totally worth learning. Most would disagree profoundly. I’ve been pushing myself to remember it, slowly but deliberately. You can check out some of the tiny experiments here. It’s not a project, more like a scratchpad that runs on opcodes and (lovely, quality) coffee.

 

                                                                 ...When in doubt, explain it to me. I'm anathema too.


 And you know what's fun? Mathematics! No, really.
I've seen it again and again: people treating floats as if they are basically the real numbers
(ℝ). They aren't!

Just take a look under the hood and you'll understand why.

"But OPQAM, I use Python/Java/Whatever. I don’t care about Assembly or floating-point registers!"

And that’s fair—until you try 0.1 + 0.2 and get 0.30000000000000004, or even 0.30000001192092895508. Then it might matter.

Can you think of a situation where such a small discrepancy could be a problem? I sure can. I can think of several, and some of them imply falling bridges.

Here’s the core issue: ints aren’t integers, and floats aren’t rationals.

 

An example

The decimal 0.1 becomes the binary: 0.0001100110011001100110011001100110011..., repeating infinitely.
But computers can’t store infinity. They cut off after some number of bits: ~23 bits for floats, ~52 bits for doubles. It’s like trying to store 1/3 in decimal — you can’t write 0.3333... forever. You round. You approximate. So do computers.

So, this means that you cannot represent 0.1 precisely.

 

The problem isn’t the mathematics — it’s representation.

How do systems and programmers deal with this?

  • Use decimal representations (decimal.Decimal) when exactness matters (e.g. money).

  • Use rational types that store fractions exactly (Fraction(1, 10)).

  • Use symbolic math when precision must be preserved throughout (SymPy, CAS software).

  • Or use fixed-point arithmetic — store cents instead of euros.

These are workarounds. The real solution? Know that floats are approximations, not truth.



                                                Accurate depiction of my two family members' faces as they read through this.


Meanwhile, back in the Real World, there was a CTF going on in the team. A colleague of mine created it and invited us all to some friendly competition for the next few weeks. Honestly? I will probably only do a couple CTFs before I turn my attention elsewhere. Having a family + hobbies + a day job does limit one's time. Still, I got to dip, and I managed to solve a challenge involving a simple—but satisfying—privilege escalation. I'm not going to give a lot of details, since it goes against the point of the 'contest' and it's actually requested by the site that we don't do it. But here's the 'trick': Classic PATH hijack. I dropped a fake ls binary into a writable dir (/tmp), positioned it early in $PATH, and executed the vulnerable binary which, instead of ls executed cat. Bang. Root shell. Dump flag. Walk away smiling.
This was, of course, allowed by purposefully using SETUID in the binary. A no-no. But there you have it. It was fun.

And yeah, sure—it’s not the most sophisticated vector ever, but the fact is, simple stuff works. You don't have to be fancy-schmancy to make something give you a 'win'. It is thus in Jiu Jitsu, and in illusionism. And so it is in hacking.

Also worth noting: a few days ago I got into a discussion with a colleague about TPM (Trusted Platform Module). I’ve blogged about the TPM issue here, but here’s the gist of it: I had to disable TPM on an old laptop (a very respectable ThinkPad x260) just to make the thing actually power off correctly. I tried everything—kernel parameters, ACPI tweaks, prayer (not really, but I totally could have!)—but nothing worked. Full shutdown always left the machine 'hot'.

Disabling TPM did the trick.

Why? Long story short: the TPM implementation on that hardware was tailored for Windows, and Linux support is... charitable at best. As for the discussion. That colleague warned me that in certain scenarios, like full power drains, I could end up with an unbootable machine if I lost TPM state. So I did what any sane people would do: I tested it again and again, simulating different scenarios.

 

Unplugged, drained, every scenario short of desoldering the CMOS battery. Result? Nothing broke. LUKS doesn’t need TPM. At least, not for the way I have this set up. In Linux, with this setting, TPM is optional unless you're deliberately tying encryption keys to it—and even then, tread carefully.

These tests were fun. They reminded me of the joy of breaking things on purpose, and the calm that comes with understanding exactly why something behaves the way it does.

So yeah, life’s been busy. I haven’t had much time to write (sentence never written by any amateur blogger ever). But I’m still here, still learning, and still hacking away.

Next up? I'll probably hit you up with some more ASM stuff as I keep on watching those videos and experimenting with stuff. Maybe some CTFs... who knows? Not me! And I'm right here. 

We’ll see.


Sunday, June 22, 2025

Playing With Bits: Of Malware Labs, Steganography and Narnia

 

 

                                    Where's mah Gibson, punk?

 

 Hi y'all!

As mentioned here, I played around with some steganography. The idea was simple and unfancy: just a .ppm file and some Python.
Why .ppm? Because it’s stupidly simple: uncompressed RGB values, no PNG compression, or arcane metadata. Just bytes (although .ppm headers can have comments, so keep that in mind).

 

The Method
                A .ppm header

 

Looking at the image above, the header is simply to get:

(0x50 0x36)  → P6: binary PPM file

(0x0a) → newline

(0x36 0x34 0x30) → 640: image width

(0x20) → space

(0x34 0x32 0x36) → 426: image height

And bam—our header info.

And then we get 3 bytes at a time, for each pixel (RGB: Red, Green, Blue), defining the color of that precise byte.

After that, pixel data: 3 bytes per pixel (RGB). Nothing else.

To hide data, I zeroed the least significant bit (LSB) of each byte, which doesn’t really alter the image in a visible way. This gives you 3 bits per pixel to encode data. Stack those bits together and slice them in blocks of 8—now you have bytes. Bytes mean ASCII. ASCII means text.

That’s the gist. You can write an entire hidden message (or image, or audio) inside another image by tampering with just the LSBs.

To make this practical, I wrote a couple of Python scripts. You can find them here. Mess with that as you see fit.

 

 

Limitations?

Sure. Quite a few:

  • .ppm headers with comments will throw things off. But you can easily code around that. Heck, have that as a homework if you'd like.
  • No encryption. Anyone with a hex editor and some free time can sniff it out. 
  • No error detection or correction. More homework.
  • Every LSB is predictably overwritten—pattern detection is trivial.
  • Any kind of compression or encryption on the container image kills the message.

 

OpSec Level: Meh

You’d probably want to:

  • Only use one LSB per pixel (maybe just Blue).

  • Randomize altered pixels with a PRNG.

  • Encrypt the payload beforehand.

     ...

But let’s be real—Chi-squared tests, histograms, stego‑detectors—if someone’s looking for it, they’ll find it.

Still, security through obscurity has its place (despite all the hate). Think authoritarian regimes where strong crypto might be illegal or where you might simply be forced to decrypt everything you own.
Just saying... it can be part of your 'onion'.


Malware Analysis Lab - The Barebones Setup

I’ve said before I’d blog about my lab setup. Never got around to it because I kept tweaking things or getting distracted by something shinier.

 Goal here: no fancy stuff, just a focus on working securely with malware.

My current setup:
- a dedicated Mini-PC.
- a managed switch which has a port isolated for the Mini-PC
- a VLAN dedicated to that host
- a Raspery Pi as Gateway/firewall/logger for that VLAN

 

A true minimal setup:

- Your laptop

- Two VMs:

  • Flare VM (Windows + analysis tools)

  • REMnux (Linux + reverse engineering toolkit)


Essentials:

  • Keep malware VM networks disabled unless strictly needed

  • Take frequent snapshots

  • Log and document everything

  • Don’t aim for perfect—just safe and functional


That's it. I won't go further into this for now. If I find the need, I'll do it later. What you want is something that will let you experiment with some safe malware samples, learn the basic tools and avoid having your home invaded by your baddies.


🧙‍♂️ Back to Narnia (OverTheWire)

Yup, I’m back at it. I paused these CTFs months ago—wasn’t getting much out of them. Realized I needed to read more about shellcode and memory handling before diving deeper.

So now I’m back. And yeah, I’m also changing my mind about avoiding walkthroughs. Everyone does them, so I might as well add my spin: fewer spoilers, more insights. You’ll find the step-by-step mess here on GitHub (don’t expect polish).

narnia0

What’s this one about? Buffer overflows and UTF-8 stuffsesses. Simple code:

                There's the Gibson <3

 

Buffer overflow by 4 bytes. The goal is to insert \xde\xad\xbe\xef in little-endian form.

First instinct was to strings, gdb, and objdump my way through. But you don’t need that. Just feed the program 20 garbage bytes (aaaaaaaa...) followed by those 4 crafted bytes.

Small snag: input is treated as UTF-8. So outputting raw 0xde didn’t work nicely.

Long story short, I learned that I couldn't easily output xde, and I did try, by looking at UTF-8/HEX lists, but it was a no go. I then decided to build a binary with printf:

printf "aaaaaaaaaaaaaaaaaaaa\xef\xbe\xad\xde" > /tmp/key
 

This solved the problem, but didn't elevate me to narnia1. A picture can explain more than a thousand words, and so a picture filled to the brim with words might be a wonder. So, here:

 

See at the top? the redirection is not returning an error message, but I am still user narnia0, and thus haven't really solved the issue. But why?
 

The issue is Linux related, to be fair. As we run the script with our payload, we're spawning a shell, but it is dying instantly since it's not connected to a tty. I needed to keep the shell opened, so I got the cat out of the bag. It forces the shell session to remain open and thus granting us permanence.
The shell is ugly, but we get elevated privileges (you can always load a nicer shell, if you really want to), and now we can search for the password. That's easy enough, so we can leave it out of this explanation.

 

That’s it for now. These blogposts are mostly memory aids and curiosity igniters. Narnia stuff goes to GitHub as I solve them, and I’ll post here only when there’s something worth rambling about.

 

Get your tools ready and keep exploring!

SUDO INIT 6: Rebooting Your C Journey

                                                                  Yeah, let's do just that. So grab some popcorn and lean back.   I'...